{"id":"05b0e07f-7cc3-4487-85b1-c08c16c79aad","arxiv_id":"2508.11709","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A conceptual model for redesigning project-based assessment around process-oriented evaluation, multi-modal evidence, AI literacy, and a dual 'GenAI Insight' evaluation lens.","lead":"University project grading is being disrupted by generative AI, because a final report can now be largely written by a chatbot. This paper proposes a new assessment framework that grades the entire project journey, including how students use AI tools, rather than only the final product.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The model's integrity guarantee rests on unverified self-reported process artifacts; without a mechanism to authenticate these traces, the conclusion's 'ensures' is unsupported.","rationale":"I read the paper in good faith as a conceptual proposal, not an empirical study. Its internal logic is generally coherent: the five redesign principles are plausible, and the worked example maps assessments to principles consistently. The use case and mapping tables (Tables III and IV) are clearly presented. However, the strongest claim uses the word 'ensures' (Section VII), which is a high bar. The single most load-bearing assumption is that the process artifacts, which are the foundation of process-oriented evaluation and GenAI Insight, are truthful. The paper repeatedly asserts that multi-modal, continuous, and supervisor-verified assessments enhance integrity, but none of these measures authenticates the student's claimed process. A student who uses GenAI to fabricate the entire journey—logs, reflections, interactions—can defeat the model. This is not a minor edge case; it is the exact adversarial scenario the paper aims to address. The reader's weakest_assumption identifies this precisely. My analysis agrees, and I recommend no change to the CONDITIONAL verdict: the model is promising but conditional on adding a mechanism to verify process artifacts, or on tempering the claim from 'ensures' to 'supports.' I did not find other independent load-bearing concerns; the cited literature appears relevant, and the argument is internally consistent once this assumption is acknowledged.","tokens_in":10740,"tokens_out":3015,"duration_ms":35278,"concrete_test":"Run a red-team experiment: N students (e.g., 20) are instructed to complete a capstone-style project while fabricating all process artifacts (Research & Planning Log, Reflection Reports, GenAI interaction documentation) using GenAI, so that the artifacts do not reflect their actual process. Independent assessors, trained on the proposed model's rubrics (Table I) and using the viva and artifact review exactly as specified, then attempt to identify which students fabricated. If detection accuracy is not significantly above chance (e.g., 50% for a binary choice), the model's integrity guarantee is falsified. Also, conduct a checklist analysis of Table II to confirm whether any assessment component is verifiable independently of student self-report.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section VII) is that the proposed PBA model with process-oriented evaluation and GenAI Insight ensures that assessments remain valid, authentic, and integrity-driven. The security of this claim depends on the veracity of student-generated process artifacts: the Research & Planning Log, Reflection Reports on GenAI use and teamwork, and documented GenAI interactions (Table II; Section IV-D). Nowhere does the paper address that these artifacts themselves can be fabricated or wholly produced by GenAI. Section VI-B argues that 'formative logs and peer feedback ... reduce opportunities for last-minute fraudulent work' and that supervisor evaluations and viva 'verify the student's ownership of their work,' but these mechanisms only verify the artifact's existence and presentation, not its truthfulness. A student can generate a week-by-week log or a reflective narrative about 'challenges overcome' using GenAI, and the viva may probe the final product, not the authenticity of every log entry. If process traces are untruthful, both the Traditional Focus and GenAI Insight lenses evaluate a fiction, collapsing the model's anti-contract-cheating logic. This is a structural, not merely empirical, gap: the model shifts trust from product to process but provides no independent verification of the process itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conceptual model for redesigning project-based assessment (PBA) in higher education in response to generative AI (GenAI). It identifies five redesign principles—multi-modal/multi-faceted assessment, AI literacy and responsible use, higher-order thinking, process-oriented evaluation, and personalised feedback—and embeds them in a 'Project Lifecycle & Assessment Hub' that is evaluated through two viewpoints: a Traditional Focus and a GenAI Insight lens. The model is instantiated in a twelve-week capstone design with nine assessment components, and the paper provides mapping tables linking these components to the redesign principles and to six PBA elements (E1–E6). The central claim, stated in Section VII, is that following the model 'ensures that assessments not only remain valid and authentic but also support the development of essential skills for future-ready graduates.'","tokens_in":2134,"tokens_out":3000,"duration_ms":69276,"significance":"The paper addresses a timely and practically important problem: how to preserve authenticity, integrity, and learning validation in project-based assessment now that GenAI can produce substantial portions of student work. The strength of the paper is its synthesis of current literature and institutional guidance (TEQSA, UNESCO, EDUCAUSE) into a structured, visually communicated model. The mapping tables (Tables III and IV) offer a concrete, ready-to-adapt template for curriculum designers, and the proposed capstone design in Table II is immediately usable. A further strength is that the claims are actionable and falsifiable: the authors explicitly state as future work the implementation and systematic evaluation of the model, which creates a clear path for empirical testing. The conceptual nature and lack of empirical validation are expected for a model-proposal paper, but the strength of the final claims goes beyond what the evidence in the manuscript can support.","major_comments":[{"comment":"The conclusion's central claim that the model 'ensures that assessments not only remain valid and authentic' is unsupported by the evidence presented. Section VI-B asserts that formative logs, supervisor evaluations, and viva 'verify the student's ownership of their work.' However, the verification mechanisms described verify the existence and presentation of process artifacts—Research & Planning Logs, reflection reports, documented GenAI interactions (Table II; Section IV-D)—but not their truthfulness. These artifacts can themselves be fabricated or wholly generated by GenAI, and a viva that probes the final product does not authenticate each log entry. The anti-contract-cheating logic therefore rests on an unverified premise: that students honestly document their process. The authors should either moderate the 'ensures' language to 'supports' or 'is designed to facilitate,' or, prefera","section":"Section VII and Section VI-B"},{"comment":"The authors state that the model 'ensures student learning and assessment security' and recommend its adoption, but the only stated validation is 'future work' — implementation and systematic evaluation of the model. The mapping tables (Tables III and IV) demonstrate alignment by construction; they are conceptual mappings, not empirical evidence that the assessments achieve the desired outcomes. For a conceptual paper, such evidence is not required, but the conclusion overreaches. A revised version should explicitly frame the model as a theoretically grounded proposal whose effectiveness requires empirical testing, and it should articulate what form that testing would take (e.g., comparative cohorts, analysis of student artifacts, instructor and student surveys).","section":"Section VII, final paragraph"},{"comment":"The 'GenAI Insight' evaluation viewpoint is not operationalized sufficiently to support the model's claim to assess AI literacy and higher-order thinking. Table I lists criteria such as 'effective prompt formulation,' 'critical evaluation of GenAI outputs,' and 'authenticity of voice,' but no rubric, rating scale, or decision rule is provided to distinguish genuine student contribution from GenAI-generated content in practice. Without such operational definitions, two supervisors could rate the same artifact very differently, undermining the model's reliability and validity—precisely the properties the paper claims to ensure. The authors should add a sample rubric or at least an annotated example showing how a GenAI Insight evaluation is conducted, including how evidence of 'guidance, curation, and significant refinement' (E4, Table I) is elicited and scored.","section":"Section V-C and Table I"}],"minor_comments":[{"comment":"The 'Weight (%)' column for 'Research & Planning Log – Formative' reads '5 and Hurdle,' which is ambiguous. The text clarifies that the weight is 5% and the assessment is a hurdle, but the table format should separate weight from hurdle status (e.g., a separate 'Hurdle' column or a note under the table).","section":"Table II"},{"comment":"There is an inconsistency in author names: the text mentions 'Pelleti et al.' for the EDUCAUSE GenAI Readiness Assessment, but reference [11] is authored by EDUCAUSE itself; later in the same paragraph 'William et al.' appears for reference [23], which is 'Williams et al.' in the bibliography. Standardize the in-text citations to match the reference list.","section":"Section III, references"},{"comment":"The phrase 'The proposed PBA Accepted in 2025 World Engineering Education Forum - Global Engineering Deans Council (WEEF-GEDC)' appears to be a stray line from the publication venue, not part of the paper's outline. It should be removed or placed in a footnote.","section":"Section I, last paragraph"},{"comment":"In the bullet list, 'The timeline distributes tasks across the semester (weeks three to 12), supporting progressive development and continuous engagement and continuous delivery' has a repetition ('continuous engagement and continuous delivery'). Reword to avoid the duplication.","section":"Section VI-A, paragraph 2"},{"comment":"The five redesign principles are presented as bullet lists of 'key points,' but the relationship between these key points and the later PBA elements (E1–E6) is not explicit. For instance, 'Focus on Higher-Order Thinking' in Section IV-C could be more directly linked to the criteria in Table I for E4 and E5. Consider adding a short mapping sentence before Table III to connect the principles to the elements, making the structure easier for readers to follow.","section":"Section IV, global"}],"recommendation":"major_revision","confidential_remarks":"The paper is conceptually useful and likely to be of interest to the journal's readership, but the gap between the model's design and its claimed assurances is significant. The issue of self-reported process artifacts being fabricated or GenAI-generated is structural to the model's integrity argument and cannot be fixed with wording alone; the authors need to either weaken the guarantees or add concrete verification mechanisms. The lack of an operationalized rubric for the GenAI Insight viewpoint also raises reproducibility concerns. I see no sign of circularity or unsupported fundamental methodology; the model is a proposal, and its claims are empirically testable. I recommend major revision rather than rejection because the central idea is defensible and the missing pieces can be supplied within the scope of a revised manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a clearly written conceptual model for redesigning project-based assessment when students can use GenAI. The genuinely useful part is the systematic mapping: five redesign principles, two evaluation viewpoints (traditional plus 'GenAI Insight'), and six PBA elements with a worked capstone design and mapping tables. That integration is the contribution; none of the individual ideas is brand new, but putting them into one operational framework is helpful for educators who need to redesign a capstone quickly.\n\nThe paper does several things well. The tables (I, III, IV) are actionable, the sample assessment grid in Table II is concrete, and the alignment with TEQSA principles gives it institutional grounding. The authors also correctly flag the risk that product-only assessment is now unreliable.\n\nThe soft spots are real but not fatal. First, the conclusion overclaims: saying the model 'ensures' validity, authenticity, and skill development is too strong for a conceptual proposal with no empirical testing. That should be softened to 'supports' or 'is designed to promote.' Second, the stress-test is right: the integrity logic depends on the truthfulness of student-produced process artifacts like the Research & Planning Log and reflection reports. A student can fabricate a weekly log or generate reflections with GenAI, and the paper doesn't address that. The viva and supervisor checks help, but they don't verify every artifact. This is a limitation the authors should acknowledge explicitly, and it's worth noting that no assessment model can fully prevent dishonesty. The point is the paper shouldn't claim the model 'ensures' integrity when a core assumption is trust in self-reporting.\n\nOne minor citation issue: Section III names 'Pelleti et al.' for the EDUCAUSE readiness assessment, but the reference is an EDUCAUSE report, not a named author. That should be cleaned up.\n\nWho is this for? Practitioners in engineering or computing education who need a defensible structure for redesigning project assessments in the GenAI era. It's not a research contribution in the empirical sense, but it is a well-organized framework that deserves serious referee attention, with revisions. I'd send it to review, and I'd probably use the mapping tables in a course redesign.","headline":"A useful practitioner framework for GenAI-era project assessment, but the 'ensures' claim and the trust-in-process-artifacts gap need work before you rely on it.","tokens_in":11456,"tokens_out":2417,"would_cite":true,"duration_ms":24058,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A five-part design keeps project-based assessment authentic when students use generative AI.","keywords":["generative AI","project-based assessment","academic integrity","process-oriented evaluation","AI literacy","higher-order thinking","personalised feedback","capstone projects"],"falsifier":"Have a cohort complete the proposed capstone design while a separate group is asked to produce all process artifacts (logs, reflections, GenAI interaction documentation) with GenAI assistance after the fact. If trained assessors cannot distinguish the fabricated artifacts from genuine ones and give equivalent grades, the model's claim to protect authenticity and integrity is falsified.","tokens_in":10637,"feed_emoji":"🎓","tokens_out":4036,"duration_ms":40567,"temperature":0.7,"pith_summary":"The paper argues that project-based assessment breaks down in the GenAI era because final artifacts can be produced or heavily shaped by AI tools, so it proposes a conceptual model that shifts evaluation to the learning process itself. The model is built on five redesign principles—multi-modal and multi-faceted assessment, AI literacy and responsible use, higher-order thinking, process-oriented evaluation, and personalised feedback—and views each project element through two lenses: traditional assessment and 'GenAI insight'. The claim is that this dual-lens, process-focused design keeps assessments valid and authentic, protects academic integrity, and develops future-ready graduate skills. A worked example of a 12-week capstone subject shows how the principles map onto concrete assessment tasks.","feed_headline":"Five principles keep capstone projects authentic in the GenAI age","feed_subtitle":"Process logs, reflections, and viva replace the final product as proof of learning.","key_machinery":"The load-bearing mechanism is the paired evaluation viewpoint: each of the six PBA elements is assessed once in a 'Traditional Focus' mode (familiar criteria such as clarity, feasibility, quality, and communication) and once in a 'GenAI Insight' mode (prompt formulation, critical evaluation of GenAI output, transparency of use, ownership of process, and reflection on ethical implications). This dual reading, together with the five redesign principles, is what the paper uses to convert GenAI from a threat to authenticity into assessed, documented learning activity.","core_discovery":"The paper's central proposal is a conceptual model in which a 'Project Lifecycle & Assessment Hub' connects five redesign principles to six elements of project-based assessment—project definition, knowledge acquisition, process management, artifact creation, communication, and reflection. Every element is evaluated from both a Traditional Focus and a GenAI Insight viewpoint: the former uses familiar criteria with minor adjustments, while the latter assesses how effectively, critically, and ethically the student engaged with GenAI, what unique human skills they demonstrated alongside it, and whether they acknowledged its use. The paper asserts that this structure ensures assessments remain va","pith_inferences":["A testable extension is to audit whether process artifacts can be fabricated: the model's integrity claim stands or falls on the truthfulness of logs and reflections, which GenAI could itself generate.","The GenAI Insight lens could be sharpened into a structured rubric for prompt engineering and output critique, making the 'critical evaluation' criterion more objective.","The model's process documentation assumes privacy-compatible collection of student–GenAI interactions; operationalising that at scale will require technical and consent infrastructure the paper does not detail.","The two-lens evaluation could be extended to program-level assessment, using the same principles to check whether a whole curriculum develops AI literacy progressively."],"forward_implications":["Educators can directly apply the five principles and the two evaluation viewpoints to existing capstone subjects, using the mapping tables as a checklist.","If the model is adopted, assessment evidence shifts from the final report to a portfolio of logs, reflections, supervisor evaluations, and viva presentations.","The worked example demonstrates that a 12-week capstone can distribute assessment across weekly checkpoints, making last-minute outsourcing harder.","The model aligns with regulatory expectations that assessment design account for both opportunities and risks of GenAI, and supports threshold-standard compliance.","Adoption would require explicit teaching of AI literacy and GenAI interaction documentation, turning responsible use into a graded outcome."],"supporting_citations":[{"why":"A regulatory report on assessment reform for the AI age supplies the guiding principles of ethical GenAI engagement and diverse, inclusive assessment design.","marker":"[1]"},{"why":"Explores the intersection of problem-based learning and AI, providing the argument that assessment strategies must incorporate AI tools while preserving authenticity.","marker":"[5]"},{"why":"Advocates the shift from product-oriented to process-oriented assessment that the proposed model takes as its core move.","marker":"[6]"},{"why":"Argues for contextual assessment design in the GenAI age, supporting the model's use of multiple viewpoints and adaptable formats.","marker":"[7]"},{"why":"Earlier work by the authors classifying assessment types in the GenAI era and outlining authentic design steps, on which the redesign principles build.","marker":"[9]"},{"why":"Proposes frameworks for integrating GenAI into educational assessment and developing AI literacy, which the paper extends to project-based contexts.","marker":"[12]"},{"why":"Provides the contract-cheating mitigation framing that motivates the process-documentation and weekly-checkpoint design.","marker":"[27]"}],"fun_headline_variants":["Process over product: new model for GenAI-era capstones","Assess the journey, not just the artifact, in the GenAI age","Dual-lens assessment: traditional + GenAI insight for projects","Redesigning project assessment for ethical GenAI use","From final product to learning process: PBA in GenAI era"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The model's integrity guarantees depend on students truthfully producing the process artifacts—research and planning logs, reflections, and documented GenAI interactions—that are used as evidence of learning; if those can be fabricated or outsourced, the assessment validates fiction rather than learning.","fun_headline_variants_meta":{"raw":{"variants":["Process over product: new model for GenAI-era capstones","Assess the journey, not just the artifact, in the GenAI age","Dual-lens assessment: traditional + GenAI insight for projects","Redesigning project assessment for ethical GenAI use","From final product to learning process: PBA in GenAI era"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000157,"raw_usage":{"total_tokens":1027,"prompt_tokens":685,"completion_tokens":342,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":429,"completion_tokens_details":{"reasoning_tokens":253}},"tokens_in":429,"tokens_out":342,"duration_ms":4094,"temperature":1.0,"reasoning_tokens":253,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:28:35.787533+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a cohort complete the proposed capstone design while a separate group is asked to produce all process artifacts (logs, reflections, GenAI interaction documentation) with GenAI assistance after the fact. If trained assessors cannot distinguish the fabricated artifacts from genuine ones and give equivalent grades, the model's claim to protect authenticity and integrity is falsified.","supporting_citations":[{"cited_title":"Assessment reform for the age of artificial intelligence,","cited_arxiv_id":null,"evidence_quote":"A regulatory report on assessment reform for the AI age supplies the guiding principles of ethical GenAI engagement and diverse, inclusive assessment design."},{"cited_title":"PBL Meets AI: Innovating Assessment in Higher Education,","cited_arxiv_id":null,"evidence_quote":"Explores the intersection of problem-based learning and AI, providing the argument that assessment strategies must incorporate AI tools while preserving authenticity."},{"cited_title":"Process not product in the written assessment,","cited_arxiv_id":null,"evidence_quote":"Advocates the shift from product-oriented to process-oriented assessment that the proposed model takes as its core move."},{"cited_title":"Contextual Assessment Design in the Age of Generative AI","cited_arxiv_id":null,"evidence_quote":"Argues for contextual assessment design in the GenAI age, supporting the model's use of multiple viewpoints and adaptable formats."},{"cited_title":"Crafting tomorrow’s evaluations: assessment design strategies in the era of generative AI,","cited_arxiv_id":null,"evidence_quote":"Earlier work by the authors classifying assessment types in the GenAI era and outlining authentic design steps, on which the redesign principles build."},{"cited_title":"Framework for adoption of generative artificial intelligence (GenAI) in education,","cited_arxiv_id":null,"evidence_quote":"Proposes frameworks for integrating GenAI into educational assessment and developing AI literacy, which the paper extends to project-based contexts."},{"cited_title":"Towards an Holistic Framework to Mitigate and Detect Contract Cheating within an Academic Institute—A Proposal,","cited_arxiv_id":null,"evidence_quote":"Provides the contract-cheating mitigation framing that motivates the process-documentation and weekly-checkpoint design."}],"review_version":1}