{"id":"1cf45976-575d-434f-b88d-31352e3ded23","arxiv_id":"2607.28179","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Twenty-five technology-enhanced tabletop exercises with 743 participants yield 24 lessons showing digital TTX platforms can scale collaborative cybersecurity training in universities.","lead":"Researchers ran 25 digital cybersecurity tabletop exercises with 743 students using an open web platform and distilled 24 practical lessons. Educators get concrete guidance on scaling team incident-response training without drowning in paper logistics.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Causal ‘increased engagement/collaboration’ and ‘scalable/replicable model’ claims rest on unbaselined self-report and single-site volume, not comparative evidence.","rationale":"The paper is correctly read as an IEEE FIE innovative-practice report: large operational footprint with an open platform and concrete lifecycle lessons is the real contribution. The reader’s weakest_assumption is exactly the load-bearing hinge—questionnaires + small instructor focus groups cannot underwrite ‘increased engagement/collaboration’ or general ‘scalable and replicable model’ without baselines or external replication. No internal inconsistency or misconduct; correctness_risk stays medium because overclaim is linguistic relative to design, not fabrication. Novelty remains moderate (extension of authors’ prior IXP work). Verdict stays CONDITIONAL: useful practitioner hypotheses and tooling for instructors who will trial the lessons, not proven learning-science gains. No stronger technical flaw (e.g., platform logic error) is evidenced in the text; the evidence gap on attribution and generalizability is the single point that, if it fails the proposed split-cohort check, forces the headline claim to be rewritten.","tokens_in":14086,"tokens_out":657,"duration_ms":26096,"concrete_test":"Re-run one repeated scenario from Table I (e.g., phishing or mobile-device loss) in a matched cohort split: half on IXP, half on the authors’ earlier non-platform format (shared docs), using the same facilitators and identical post-questionnaire items plus platform-exportable process metrics (milestone reach rates, email thread volume/time). If self-reported engagement/collaboration and process metrics show no reliable advantage for IXP after novelty/order controls, the Abstract/§VI causal and ‘demonstrated model’ wording should be softened to operational lessons only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (Abstract, §I, §VI) is that IXP-delivered digital TTXs are a demonstrated scalable and replicable model that increased engagement and collaboration across 25 runs / 743 participants, with clearer instructor visibility and reduced workload, plus 24 transferable lessons. What must be true: (1) the observed gains are attributable to the technology-enhanced format rather than novelty, course credit, competition selection, or facilitator skill; (2) volume and lessons generalize beyond the authors’ institution, self-designed scenarios, and instructor pool. Section V opening states the lessons come from post-exercise trainee questionnaires and focus groups with 8 instructors total; Table I shows heavy concentration in the authors’ own cybersecurity courses (repeated scenario families) plus selected extracurriculars. There is no baseline arm (pen-and-paper or SharePoint as in their prior [13]), no validated engagement/collaboration instruments, no pre/post learning measures, and no independent external delivery sites reported. Platform logs (milestones, emails) support process visibility claims but are not used as comparative outcome evidence for ‘increased’ engagement. Thus the demonstration language outruns the design: operational feasibility at scale under author facilitation is shown; causal improvement and broad replicability remain assumptions.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This innovative-practice paper reports on integrating technology-enhanced cybersecurity tabletop exercises (TTXs) via the open-source INJECT Exercise Platform (IXP). The authors describe the INJECT Process (understanding–specification–preparation–execution–reflection), platform capabilities (milestone-driven injects, simulated tools/email, dashboards, YAML/editor authoring, on-demand and multi-tenant runs), and 25 deliveries (2024–2026) totaling 743 participants across courses and extracurricular events (Table I). From post-exercise trainee questionnaires and instructor focus groups (8 instructors), they distill 24 lessons spanning the lifecycle and argue that digital TTXs are a scalable, replicable model that increases engagement/collaboration, reduces instructor workload, and improves visibility into team decision-making.","tokens_in":14376,"tokens_out":1257,"duration_ms":31122,"significance":"For computing/cybersecurity education practice, the contribution is substantial and usable: a documented multi-year deployment at non-trivial scale, an open platform with exercise definitions and tooling, and concrete facilitation/design lessons organized by lifecycle phase. Strengths that should be credited include open-source IXP and exercise library, explicit milestone/tool design guidance, real-time instructor views, and exportable logs that enable research reuse. Even if causal effectiveness claims are tempered, the operational feasibility evidence and practitioner guidance are valuable for FIE-style innovative practice and for curriculum designers adopting digital TTXs.","major_comments":[{"comment":"Abstract, §I, and §VI claim the practice “increased engagement and collaboration,” yielded “actionable insight into student learning,” and that the authors “demonstrate” a “scalable and replicable model.” Section V states the evidence base is post-exercise trainee questionnaires plus focus groups with 8 instructors; Table I shows concentration in the authors’ own repeated course scenarios plus selected extracurriculars. There is no baseline arm (pen-and-paper/SharePoint as in prior work [13]), no validated engagement/collaboration instruments, and no pre/post learning measures. Platform logs support process-visibility claims but are not used comparatively for “increased” outcomes. The load-bearing causal and generalizability language should be revised to match the design: demonstrated operational feasibility and instructor-perceived benefits under author facilitation, with transfer treat","section":"Abstract, §I, §V opening, §VI"},{"comment":"The replicability claim (Abstract/§VI) rests on volume plus lessons, but external independent delivery sites, non-author facilitators running full scenarios without the design team, and non-cyber domains are not reported. Table I’s repeated scenario families and single primary institution make “scalable and replicable model for … others requiring team-based problem-solving” stronger than the evidence. Either report external reuse data if available, or narrow the claim to “a scalable delivery model in our setting, with lessons intended to transfer,” and state boundary conditions (facilitator skill, scenario quality, institutional context) explicitly in §VI.","section":"Abstract, Table I, §VI"},{"comment":"§V.E and the conclusions correctly emphasize that reflection is where learning happens and that on-demand automation trades away structured debrief. Given that stance, the paper’s own outcome claims still lean on immediate post-exercise self-report rather than structured reflection products (action commitments, decision comparisons tied to milestones, delayed follow-up). Strengthening the manuscript does not require a new RCT, but it does require aligning claims with what was measured—or briefly reporting any debrief artifacts/scores actually used—so the “learning” language is not carried only by engagement impressions.","section":"§V.E, §VI"}],"minor_comments":[{"comment":"Table I header glyphs (discussion vs simulation) are hard to parse in plain text/print; add an explicit column legend and spell out exercise type in the table body.","section":"Table I"},{"comment":"DOI is still “TODO”; fix before camera-ready. Several URLs are dated 2026 access—ensure consistency with the proceedings timeline.","section":"Title block / References"},{"comment":"The paper says “24 lessons” but the enumerated lettered items under §V are easy to miscount; add a compact numbered inventory (or appendix checklist) mapping each lesson to a phase for reuse by instructors.","section":"§V"},{"comment":"Figures 2–6 are helpful but depend on screenshots; ensure captions stand alone (what milestone clustering distance means for a practitioner; what an instructor should do when a team is an outlier in Fig. 5).","section":"§V.C–§V.E"},{"comment":"Related work is current; still, briefly contrast IXP’s milestone///tool model with Watkins et al.’s AI-inject proposal on what is actually deployed vs. proposed, to sharpen novelty for readers.","section":"§II.B"},{"comment":"Minor copyediting: spacing anomalies (“first -time”, “human -readable”, “Y AML”), and consistent expansion of TTX/IXP on first use in each major section.","section":"Throughout"}],"recommendation":"minor_revision","confidential_remarks":"Appropriate as FIE innovative-practice: the operational corpus is real and the lessons are usable. Main editorial risk is Abstract/conclusion overclaim relative to unbaselined self-report; requiring tone-down plus explicit limitations should be sufficient—no need for major new experiments for this venue. Heavy self-citation of IXP prior work is expected for a platform evolution report and is not disqualifying."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a useful innovative-practice report, not a learning-science result. What is new is volume and packaging—25 IXP runs, 743 participants across courses and extracurriculars, a named five-phase INJECT Process, and 24 concrete lessons—not the platform itself (already in their ITiCSE 2024 work).\n\nWhat they do well is operational honesty about delivery. Table I is clear. The lessons track the lifecycle (audience fit, milestone logic as the hard cognitive shift, team size of three, instructor templates, real-time dashboards, on-demand vs facilitated tradeoffs, reflection as where learning sticks). The open tooling, YAML/editor path, and log/export story make the work reproducible for someone who wants to try it. Self-citation is heavy but expected for a platform evolution paper; circularity is low.\n\nSoft spot, in proportion: Abstract/§I/§VI say they “observed increased engagement and collaboration” and “demonstrate” a scalable replicable model. Evidence is post-exercise questionnaires plus focus groups with eight instructors, mostly their own courses and repeated scenario families, no baseline (pen-and-paper or SharePoint), no validated instruments, no pre/post learning measures, no independent external sites. Platform logs support visibility and process claims; they do not establish causal gains over traditional TTXs. Stress-test is right on the overclaim; wrong if read as “the whole paper fails”—feasibility at author-facilitated scale is shown; transfer and causal improvement are not.\n\nWho it’s for: instructors and curriculum designers who will treat the lessons as hypotheses to try, and platform builders. Not for someone hunting controlled effect sizes. Math/data/citations look fine for the genre. I’d send it to peer review at FIE; I’d cite the lessons and scale numbers if I were running digital TTXs, with a caveat on the engagement claim. Engage if you teach incident response or build exercise tooling; skip if you need comparative learning outcomes.","headline":"Solid FIE practice paper: real scale (25 runs, 743 people) and usable lessons on an open TTX platform; causal “increased engagement” language outruns the evidence.","tokens_in":15019,"tokens_out":510,"would_cite":true,"duration_ms":16733,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A web platform that automates cybersecurity tabletop exercises can scale team training in universities, backed by 25 runs and 24 lessons from 743 participants.","keywords":["collaborative learning","cybersecurity education","tabletop exercise","TTX","incident response","simulation-based learning","learning analytics","INJECT Exercise Platform"],"falsifier":"Run the same scenarios for matched cohorts with and without the platform (or with paper TTXs), using pre/post skill measures and blinded ratings of team decisions; if engagement, collaboration quality, and learning gains do not differ, the central scalability-and-benefit claim fails.","tokens_in":14973,"feed_emoji":"🛡️","tokens_out":847,"duration_ms":20139,"temperature":0.7,"pith_summary":"Professional cybersecurity tabletop exercises train teams to coordinate during incidents, but universities rarely use them. This paper shows how a purpose-built web platform can automate scenario delivery, log team actions, and support assessment so the format fits ordinary courses. Across 25 exercises with 743 students and competition finalists, the authors report higher engagement and collaboration, lower instructor overhead, and clearer visibility into how teams navigate scenarios. They organize the work as a five-phase lifecycle and publish 24 concrete lessons on audience design, milestone logic, team size, facilitation, and post-exercise reflection. The claim is that digital tabletop exercises are a scalable, reusable model for cybersecurity classes and other team problem-solving courses.","feed_headline":"Digital tabletops scale cyber team training in class","feed_subtitle":"25 exercises, 743 students, and 24 lessons show how automation fits university courses.","key_machinery":"The INJECT Exercise Platform (IXP) plus the INJECT Process: a web environment that drives scenarios via timed or milestone-triggered injects, simulated tools and email, and logged actions, structured across understanding, specification, preparation, execution, and reflection phases.","core_discovery":"Technology-enhanced tabletop exercises delivered through an automated web platform are a scalable and replicable model for university cybersecurity education: automating inject delivery, milestone-driven scenario flow, and interaction logging reduces instructor workload, raises realism, and yields actionable insight into team decision-making, as shown across 25 exercises with 743 participants and distilled into 24 lifecycle lessons.","pith_inferences":["The bottleneck will shift from delivery logistics to scenario design skill—especially milestone logic—so faculty development may matter more than more platform features.","If AI-assisted scoring of free-text replies matures, instructor-in-the-loop evaluation could scale without forcing every exercise into multiple-choice form.","Institutions that treat scenario libraries as shared curriculum assets, not one-off events, will capture most of the claimed reuse benefit.","Comparative studies against other active-learning formats (not only paper TTXs) would clarify whether the gains are format-specific or mainly from structured teamwork time."],"forward_implications":["Instructors can reuse scenario definitions across cohorts with little rework once milestone logic and content are stable.","Real-time milestone dashboards let facilitators spot stuck teams during a run instead of waiting for paper debriefs.","On-demand fully automated exercises can reach large enrollments, at the cost of less flexible free-form assessment and live facilitation.","Structured reflection that ties logged decisions to next-step commitments becomes the main lever for lasting behavior change.","The same digital TTX pattern can transfer to other team-based problem-solving courses beyond cybersecurity."],"fun_headline_variants":["Automated TTXs scale cyber team training across 25 campus exercises","Web platform turns tabletop drills into measurable cyber class practice","25 exercises, 743 students: automated injects cut instructor TTX load","IXP-delivered tabletops make campus cyber incident response repeatable","Milestone-driven digital TTXs yield 24 lessons for cyber curricula"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That post-exercise questionnaires and small instructor focus groups are enough to show that engagement, collaboration, and learning improved because of the digital format, and that those gains will hold outside the authors’ courses and scenarios.","fun_headline_variants_meta":{"raw":{"variants":["Automated TTXs scale cyber team training across 25 campus exercises","Web platform turns tabletop drills into measurable cyber class practice","25 exercises, 743 students: automated injects cut instructor TTX load","IXP-delivered tabletops make campus cyber incident response repeatable","Milestone-driven digital TTXs yield 24 lessons for cyber curricula"]},"model":"grok-4.5","effort":"low","cost_usd":0.003602,"raw_usage":{"total_tokens":1156,"prompt_tokens":779,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":36024000,"prompt_tokens_details":{"text_tokens":779,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":285,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":779,"tokens_out":92,"duration_ms":6102,"temperature":1.0,"reasoning_tokens":285,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T15:28:28.877229+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same scenarios for matched cohorts with and without the platform (or with paper TTXs), using pre/post skill measures and blinded ratings of team decisions; if engagement, collaboration quality, and learning gains do not differ, the central scalability-and-benefit claim fails.","supporting_citations":[],"review_version":1}