{"id":"1d9bb142-da66-4a33-bd6d-7f9b7b5c46ff","arxiv_id":"2501.02063","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors assemble a multi-format corpus of 80 institutional generative-AI guidelines and run descriptive keyword analyses across six continents.","lead":"This paper describes AGGA, a new collection of 80 documents on how universities and other institutions regulate the use of generative AI tools. It is intended as a benchmark for natural language processing studies of academic AI guidelines.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Supplementary Table 1 lists non-university and non-guideline items (U.S. Dept.","rationale":"The reader's weakest assumption correctly identifies the load-bearing premise: the dataset is only valuable as a corpus of academic guidelines if the 80 items are genuinely official university-issued guidelines. The paper's own Supplementary Table 1 undermines this premise. Several rows are national government documents or sector-level frameworks, and several others are not guidelines at all, which makes the claim '80 academic guidelines collected from official university websites' factually inaccurate as stated. This is not a stylistic or labeling quibble; it affects which documents the NLP analyses are performed over, and the paper's technical validation is intended to demonstrate inclusiveness of academic guidelines. The concern is concrete and verifiable, and the evidence is internal to the manuscript. I do not see another concern with greater leverage on the central claim. At the same time, the resource is not beyond repair: many entries are genuine university guidelines, the dataset is openly archived with a persistent DOI, the code is available, and the curation issues are correctable. The reader's CONDITIONAL verdict is therefore appropriate; my stress-test does not move it. A rejection would be too harsh given the salvageable core, and an acceptance would ignore the mismatch between the dataset's scope and its contents.","tokens_in":13550,"tokens_out":3168,"duration_ms":32023,"concrete_test":"Re-run the dataset curation audit with a pre-registered binary rubric: an item counts as an academic guideline iff (i) the issuing body is the university or an official university organ, not a ministry, department of education, sector association, or third-party lab; (ii) the document is a policy, guidance, or statement on GAI/LLM use in academic or educational settings; and (iii) the URL resolves to that body's official page, not a local file path or third-party blog. Independently classify all 80 rows using the listed URLs and references, and check for duplicate documents. Then recompute the continental frequency and network analyses on the compliant subset. If the compliant count is materially below 80, the Abstract's scope claim and the 'six continents' validation need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that AGGA comprises 80 academic guidelines for GAI/LLM use, 'meticulously collected from official university websites.' The dataset's own Table 1 contradicts this scope. Rows 30 (U.S. Department of Education report), 36 (Peru's National AI Strategy), 37 (Chile's Ministry of Science guidelines), 58 (China's MOST Responsible Research Conduct), 59 (Australian Framework for AI in Schools), and 80 (Russell Group principles) are not university academic guidelines; they are national, governmental, or multi-institution documents. Other rows are not guidelines at all: row 47 is a Tsinghua forum page, row 42 is an editorial policy, row 43 is a report on judicial-sector readiness, and row 32 is guidance for ChatGPT use in the justice sector. Rows 51 and 54 duplicate the same CUHK student guide under different university names, while rows 46, 56, and 75 contain naming/typing errors ('Weseda', 'SUM', 'Arificial Intelligence'). Because the dataset's stated value as a benchmark depends on every item being an authentic university-issued academic guideline, these inclusions overstate the corpus. The word-frequency and network analyses are computed over this mixed set, so the descriptive conclusions inherit the composition error. The fix is tractable: re-audit each row and report the subset that genuinely satisfies the stated inclusion criteria.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AGGA, a dataset claimed to contain 80 academic guidelines for the use of Generative AIs and Large Language Models, collected from official university websites across six continents. It describes the collection and exclusion criteria, the dataset's file formats (DOCX, PDF, XLSX) on Harvard Dataverse under a CC0 license, and two descriptive analyses: continental keyword-frequency charts and a keyword-continent network graph. The authors position the dataset as a resource for NLP and requirements-engineering tasks such as ambiguity detection, categorization, and structure assessment, and they provide usage notes and a GitHub repository with processing code.","tokens_in":13908,"tokens_out":4822,"duration_ms":43517,"significance":"If the dataset matched its description, AGGA would be a useful open resource for studying institutional AI-use policies, with the practical merit of being openly licensed, multi-format, and geographically broad. The authors also provide a data DOI and a code repository, which are appropriate for a dataset paper. However, the central value of the corpus depends on each item being an authentic, university-issued academic guideline, and the manuscript's own table contradicts that premise: several rows are national strategies, government reports, multi-institution principles, or non-guideline web pages. Since the word-frequency and network analyses are computed over this mixed collection, the descriptive conclusions inherit the composition error. The dataset is potentially salvageable through a systematic re-audit and re-release, but the current version overstates its scope.","major_comments":[{"comment":"Rows 30, 36, 37, 58, 59, and 80 are not university-issued academic guidelines: they are a U.S. Department of Education report, Peru's national AI strategy, Chile's Ministry of Science guidelines, China's MOST research-conduct guidelines, the Australian Framework for Generative AI in Schools, and the Russell Group principles. Row 10 is also not a university but the African Observatory on Responsible AI. These entries directly contradict the abstract and Methods, which define AGGA as '80 academic guidelines' collected from 'official university websites' and describe excluding universities without official guidelines. The '80' count is therefore misleading, and all quantitative descriptions of the corpus (188,674 words, per-continent statistics) are computed over an out-of-scope set. A complete re-audit is needed, with either removal of non-conforming items or a revised, explicit statement of broader inclusion criteria, followed by recomputation of corpus statistics.","section":"Supplementary Table 1"},{"comment":"Several entries are not academic guidelines at all: row 32 is guidance for ChatGPT use in the justice sector, row 42 is an editorial policy of a university press, row 43 is a report on judicial-sector readiness in Latin America, row 47 is a Tsinghua forum page, row 73 is a research-group page on artificial intelligence at the University of Padua, and row 28 is the Montreal Declaration on Responsible AI rather than a University of Montreal guideline. Labeling these documents as 'academic guidelines' overstates the dataset's scope and can distort the continental keyword profiles, such as the 'justice' emphasis reported for South America. The authors should either remove such non-guideline entries or clearly re-scope the dataset's title and description.","section":"Supplementary Table 1, rows 32, 42, 43, 47, 73, 28"},{"comment":"Provenance errors undermine the claim of meticulous collection from official websites. Row 51 is labeled University of Hong Kong but links to a CUHK document; row 54 lists the same CUHK guide under The Chinese University of Hong Kong, making a duplicate or misattribution; row 53 (National Taiwan University) reuses reference [60], which points to National Tsing Hua University; reference [10] gives 'https://chat.openai.com/chat' rather than the Cairo University policy document; and reference [86] is a local file path on the author's computer. Each row should be verified against the named institution's official web presence, and every reference should be replaced with a stable, publicly resolvable URL or DOI.","section":"Supplementary Table 1, rows 51, 53, 54; references [10], [60], [86]"},{"comment":"The keyword-frequency and network analyses are descriptive summaries of the dataset as currently composed. Because the dataset includes government documents, multi-institution frameworks, and non-guideline web pages, the reported continental themes (e.g., 'justice' for South America, 'frameworks' for Oceania) cannot be attributed to university academic guidelines. If the dataset composition is corrected, these analyses and visualizations must be regenerated on the verified subset; as presented, they do not validate the inclusiveness of a corpus of university guidelines specifically.","section":"Technical Validation, Figures 2 and 3"}],"minor_comments":[{"comment":"There are several typos and naming errors that should be fixed: row 22 'Gudelines' should be 'Guidelines'; row 46 'Weseda' and 'Arificial' should be 'Waseda' and 'Artificial'; row 48 'SUM' should be 'SMU'; and row 51 'Nayang' and 'NUT' should be 'Nanyang' and 'NTU'.","section":"Supplementary Table 1, rows 22, 46, 48, 51"},{"comment":"Figure 1 lists 'Applying an XML Schema (XSD) to standardize the documents' as a processing stage, but the Data Records section describes only DOCX, PDF, and XLSX files and never mentions an XML instance or schema. Clarify whether an XML version exists or remove that step from the workflow diagram.","section":"Figure 1 and Data Records"},{"comment":"The text says the dataset is licensed under CC0 1.0 'provided proper credit is given,' which is inconsistent with CC0, as CC0 does not require attribution. Rephrase to say that attribution is appreciated but not required, or state the actual license terms precisely.","section":"Data Records, licensing sentence"},{"comment":"The caption 'Analytics Workflow of the analysis of GAI/LLM guidelines' does not match the table's content, which lists the 80 guideline entries rather than an analytics workflow. Rename the caption to describe the dataset inventory or the selection results.","section":"Table 1 caption"},{"comment":"Institution names are inconsistently capitalized and formatted (e.g., 'Universidad de san andres' and 'pontificia universidad católica de chile'). Standardize institution names and country spellings to avoid ambiguity in downstream use of the Excel file.","section":"Supplementary Table 1, rows 35 and 46"}],"recommendation":"major_revision","confidential_remarks":"The core idea of a curated, open dataset of university AI guidelines is worthwhile, and the authors have done the community a service by sharing the corpus and code. However, the current submission cannot be accepted as a data descriptor because its central claim—that all 80 items are university-issued academic guidelines collected from official websites—is contradicted by the manuscript's own table. The fix is tractable: re-audit every row, restrict or explicitly re-scope the inclusion criteria, correct the duplicate and misattributed entries, and regenerate the descriptive analyses. I would also gently flag that references [4] and [5] are close companion papers by the same authors describing the same or overlapping corpora; the relationship should be clarified to avoid self-citation or duplicate-publication concerns. After a thorough revision, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: AGGA is a genuinely useful dataset, but it is not what its abstract says it is. The 80 rows in Supplementary Table 1 are not all academic guidelines from university websites; a non-trivial fraction are government strategies, ministry guidance, an observatory report, a forum page, and a multi-university network statement. That mismatch is the whole ballgame for a data descriptor, and it is fixable.\n\nWhat's new and good: the corpus itself—80 (or, after cleaning, maybe 70–75) documents in DOCX, PDF, and XLSX, CC0-licensed on Dataverse, with a companion GitHub notebook. That is a real contribution for requirements engineering and AI-policy text mining; there is no equivalent open corpus of this scope. The descriptive keyword and network analyses are not deep, but they are honest descriptive summaries and the code is straightforward to rerun. The paper is transparent about its own stopword choices and processing steps.\n\nSoft spots: the inclusion criteria are stated in the Methods ('official guidelines specifically addressing the use of GAIs... from official university websites'), but Table 1 contradicts them. Rows 30, 36, 37, 58, 59, and 80 are national or multi-institution documents; row 42 is an editorial policy; row 43 is a report on judicial-sector AI readiness; row 47 is a conference forum page; row 28 is the Montreal Declaration, not a university guideline. There are also curation slip-ups: row 51 lists 'Nayang' (Nanyang) and 'NUT' (NTU), row 46 'Weseda' (Waseda) and 'Arificial', row 48 'SUM' (SMU); rows 43 and 54 appear to point to the same CUHK PDF, and rows 52–53 share a reference incorrectly. These are the kind of errors that make a benchmark unreliable if someone uses it naively.\n\nThe keyword and network analyses are computed over this mixed set, so the continent-level comparisons inherit the composition problem. That is a real issue, but it is not fatal: the descriptive conclusions are soft and would probably shift modestly after re-auditing.\n\nCitation pattern: the two companion papers (2406.18842, 2501.00957) are the same authors' related work; citing them is fine, but you should know this paper is one of a cluster.\n\nWho this is for: researchers building or evaluating corpora of AI-use policies; anyone doing NLP on institutional guidelines. It deserves peer review, but as a major revision: re-audit every row, correct the duplicates and typos, and either tighten the scope claim or rename the corpus to 'policy documents and guidelines.' I would engage with it after that.","headline":"A useful, openly archived dataset whose central scope claim is undercut by its own table; fixable with a careful re-audit.","tokens_in":14310,"tokens_out":2926,"would_cite":false,"duration_ms":26968,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The AGGA dataset packages 80 official university guidelines on generative AI and large language models into an openly licensed, multi-format corpus built for NLP research.","keywords":["generative AI","large language models","academic guidelines","NLP benchmark","requirements engineering","text mining","higher education policy","dataset"],"falsifier":"Open Table 1 and count how many of the 80 listed sources are issued directly by a university's own central administration, faculty, or department; any substantial number of entries that turn out to be national strategies, ministry reports, or consortium statements would falsify the paper's description of 'academic guidelines collected from official university websites' and would shrink the effective scope of the corpus.","tokens_in":13350,"feed_emoji":"📚","tokens_out":8021,"duration_ms":72315,"temperature":0.7,"pith_summary":"AGGA is a curated collection of 80 official university guidelines on the use of generative AI and large language models in academic settings, drawn from institutions across six continents and totaling 188,674 words. The paper's central claim is that this openly licensed, multi-format corpus gives NLP researchers a reproducible foundation for tasks common in requirements engineering, such as synthesizing policy models, detecting ambiguity, categorizing requirements, and assessing document structure. The authors also argue the dataset can be annotated into a benchmark for comparing how different institutions regulate AI use. If the claim holds, AGGA fills a documented gap: fewer than 10 percent of schools and universities have formal AI-use policies, and scattered proprietary documents have made experiments hard to replicate.","feed_headline":"80 university AI policies become one open dataset","feed_subtitle":"Six continents, 188,674 words, and a public-domain license give NLP researchers a shared testbed for AI policy text.","key_machinery":"The central object is the AGGA dataset, a structured compilation in which each guideline is tagged by continent, country, university, document name, and page count, with full text stored in DOCX, PDF, and XLSX files. The supporting machinery is a text-processing pipeline that tokenizes, removes stopwords, stems, and lemmatizes the documents, then runs frequency and network analyses to display continent-level keyword patterns. This pipeline demonstrates the dataset's usefulness for text mining and serves as a template for downstream NLP tasks.","core_discovery":"The core discovery is the AGGA dataset itself: 80 academic guidelines for generative AI and large language models, collected from official university websites, normalized into a consistent structure, and released in Word, PDF, and Excel formats under a public-domain license. The paper reports that the corpus contains 188,674 words and spans six continents, and it validates the collection through text-mining analyses that show distinct regional keyword emphases. The intended contribution is a shared resource that lets researchers apply NLP methods to academic AI policy, with the dataset structured so that model synthesis, abstraction identification, ambiguity detection, and requirements categorization can be performed and compared across institutions.","pith_inferences":["The dataset could be turned into a longitudinal resource: re-crawling the same universities' pages at later dates would reveal how AI-use policies change as the technology and regulatory climate evolve.","A user should filter the 80 entries by issuing body before benchmarking; several rows are national strategies, ministry reports, or consortium principles rather than university-issued academic guidelines, so a stricter university-only subset may give different results.","The continent-level keyword differences suggest that an NLP model trained on one region's guidelines may transfer less cleanly to another; a cross-validation experiment could test that directly.","Pairing AGGA with a comparable corpus of industry or government AI guidelines would allow direct comparison of how academic and non-academic sectors frame responsible AI use."],"forward_implications":["Researchers can use AGGA as a shared benchmark for ambiguity detection, requirements categorization, and equivalent-requirement identification in AI policy documents.","The multi-format release lets one team do close qualitative reading in Word and PDF while another runs quantitative scraping and modeling on the structured table.","Because the corpus spans six continents, comparative studies of regional differences in AI governance language become possible with a single accessible dataset.","The public-domain license means institutions can reuse or adapt the guidelines, and researchers can redistribute derived annotations without restriction.","The planned expansion on an open code repository lets the corpus track newly issued university policies as they appear."],"supporting_citations":[{"why":"The dataset's own archived record, supplying the persistent identifier, version, and public-domain license.","marker":"[8]"},{"why":"A global survey establishing that fewer than 10 percent of schools and universities have formal AI-use policies, the gap AGGA addresses.","marker":"[2]"},{"why":"A compilation of university generative-AI guidelines from around the world that informs the collection and framing.","marker":"[1]"},{"why":"Companion analysis by the same authors of the global landscape of academic GAI and LLM guidelines, which provides the analytic context for the dataset.","marker":"[4]"}],"fun_headline_variants":["80 unis, 6 continents, 188k words of AI policy","AGGA dataset: 80 university AI guidelines in public domain","One dataset to study all academic AI policies? 80 included","Open data: 80 university rules on generative AI, ready for NLP"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's value depends on the assumption that all 80 documents are truly official academic guidelines issued by universities, not government reports, ministry strategies, or third-party principles.","fun_headline_variants_meta":{"raw":{"variants":["80 unis, 6 continents, 188k words of AI policy","AGGA dataset: 80 university AI guidelines in public domain","One dataset to study all academic AI policies? 80 included","Open data: 80 university rules on generative AI, ready for NLP"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000309,"raw_usage":{"total_tokens":1705,"prompt_tokens":827,"completion_tokens":878,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":802}},"tokens_in":443,"tokens_out":878,"duration_ms":8854,"temperature":1.0,"reasoning_tokens":802,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:13:32.459875+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Open Table 1 and count how many of the 80 listed sources are issued directly by a university's own central administration, faculty, or department; any substantial number of entries that turn out to be national strategies, ministry reports, or consortium statements would falsify the paper's description of 'academic guidelines collected from official university websites' and would shrink the effective scope of the corpus.","supporting_citations":[{"cited_title":"AGGA: A Dataset of Academic Guidelines for Generative AIs,","cited_arxiv_id":null,"evidence_quote":"The dataset's own archived record, supplying the persistent identifier, version, and public-domain license."},{"cited_title":"The global landscape of academic guidelines for generative AI and Large Language Models","cited_arxiv_id":"2406.18842","evidence_quote":"Companion analysis by the same authors of the global landscape of academic GAI and LLM guidelines, which provides the analytic context for the dataset."}],"review_version":1}