{"id":"3c482401-ea45-4493-8c4b-e18615b21ff1","arxiv_id":"2412.04185","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using course materials as context, GPT-4 Turbo can generate structurally annotated quiz questions, but it frequently produced relational semantic annotations incorrectly and about one in three questions had factual errors, so significant human review is still required.","lead":"This paper tests whether GPT-4 Turbo, fed with course materials through retrieval-augmented generation, can write quiz questions with the semantic markup needed by an adaptive learning system. It finds that the model handles simple structural markup well but fails at linking questions to course concepts, and roughly a third of its questions contain content errors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing weak point is the unreported expert-rater reliability behind the paper's counts of question fit, solvability, and content errors.","rationale":"I agree with the reader that the weakest assumption is evaluation reliability. The paper is honest and its limitations section acknowledges missing student validation and control items, but it does not acknowledge or remedy the absent inter-rater reliability. The dangerous wrong-question example is persuasive evidence that human review is needed, so I would not reject the paper; however, the exact counts and the 'often' generalization cannot be audited without rater-level data. Since the reader already assigned CONDITIONAL, my stress test does not move the verdict. No internal inconsistency or unsupported formal claim was found; the concern is about the strength of empirical support for the central claim.","tokens_in":16210,"tokens_out":6430,"duration_ms":68866,"concrete_test":"Obtain the anonymized rating matrix (experts × 30 generated questions × six Likert items, difficulty, and content-error flag) from the authors, or re-run the evaluation with at least two independent, non-author experts on the same 30 questions. Compute per-item inter-rater agreement (weighted Cohen's kappa or Krippendorff's alpha) and report the number of raters per item. If agreement is below 0.6 or only one rater rated each item, re-report the counts as ranges and downgrade 'often' to pilot-level qualitative language; if agreement is high with at least two raters per item, the current counts stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative findings — 28/30 fit, 27/30 solvable, 11/30 with content errors, and the qualitative claim that questions 'often did not meet educational standards' — are all derived from an expert survey whose design is described in §4.5.1, but the paper never states how many experts participated, whether each generated question was rated by more than one expert, or any inter-rater agreement statistic. With a single rater per item, every count becomes a proxy for that one rater's severity and expertise; with multiple raters, unreported disagreement could hide unstable judgments. The same measurement-validity gap affects the conclusion that relational annotations are poor, which is reported without per-question annotation accuracy counts. Section 5.4 explicitly concedes the absence of student validation and control items, so there is no anchor for rating severity. This makes the empirical support for the central claim rest on an unsubstantiated assumption about the reliability of the expert judgments.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether GPT-4 Turbo, combined with retrieval-augmented generation and the sTeX semantic markup framework, can generate course-specific, semantically annotated quiz questions for an adaptive learning assistant. The authors describe a prompt-based pipeline, generate 30 questions across six topics of an 'Artificial Intelligence I' course, and evaluate them through an expert survey. They report that structural semantic annotations are generated reliably, while relational annotations are not; that 28/30 questions fit the teaching material, 27/30 are solvable, and 11/30 contain content errors; and that question quality often falls below educational standards, requiring human intervention.","tokens_in":16318,"tokens_out":2947,"duration_ms":32498,"significance":"If the empirical findings are reliable, the paper makes a useful contribution to automated question generation in higher education by showing both a feasible direction (structural semantic annotation with RAG) and a clear limitation (relational concept linking and autonomous question quality). The authors are honest about several threats to validity, including the absence of student validation and control items, and they provide a detailed account of prompt iteration and the sTeX infrastructure. The main quantitative claims, however, rest on an expert evaluation whose rater design and reliability are not reported, so the strength of the evidence is currently indeterminate.","major_comments":[{"comment":"The paper's central quantitative results (28/30 fit, 27/30 solvable, 11/30 with content errors) rest entirely on the expert survey, but the survey design is under-specified: the manuscript does not state how many experts participated, whether each generated question was rated by more than one expert, or any inter-rater agreement statistic. Since the ratings are subjective Likert judgments plus a free-text error field, the reported counts cannot be interpreted without reliability information. The limitation statement in §5.4 explicitly notes the absence of control items and student validation, leaving no anchor for rating severity. Please report the number of raters and per-item agreement, or reframe the counts explicitly as exploratory single-rater observations that do not support generalization.","section":"§4.5.1 and §5.1"},{"comment":"The conclusion that relational annotations 'exhibited poor integration' is supported only by qualitative observation and a failed function-calling/RAG augmentation; no quantitative measure of relational annotation correctness is provided, such as the fraction of generated symbol references that resolve to existing sTeX symbols, the number of incorrect module imports, or precision/recall of concept links. Because RQ2 asks 'to what extent' LLMs can annotate questions semantically, the headline claim about relational annotations needs a metric to be assessed. Even a coarse count (e.g., correct vs. incorrect |symref or |sr usages) would make the finding reproducible.","section":"§5.3 and RQ2"},{"comment":"The 11/30 content-error count is not accompanied by a coding scheme or error taxonomy, nor by a per-topic breakdown (only arc consistency and propositional logic are named qualitatively). Without a predefined definition of what counts as a 'content error' and evidence of rater agreement on that definition, the count is not independently verifiable. This matters because the paper's educational-standards conclusion leans heavily on this number; please provide the coding criteria and, if feasible, the distribution of errors across the six topics.","section":"§5.2"}],"minor_comments":[{"comment":"There is a typo: 'GTP-4-Turbo' should be 'GPT-4-Turbo'.","section":"§4.2"},{"comment":"The phrase 'our main objective in AGQ' uses the wrong acronym; it should be 'AQG' (automated question generation) as used elsewhere.","section":"§4.5.2"},{"comment":"The illustrative natural-deduction question is described as 'generated in an earlier experiment'; please clarify whether it is part of the 30-question evaluation set or an additional qualitative example, since it is used to support a quantitative claim.","section":"§5.2"},{"comment":"The full prompt and the 30 generated questions are not included in the manuscript or a supplement, and §4.2 says no public pipeline instance is available; making these available in supplementary material would substantially improve reproducibility.","section":"§4.2 and §4.3"},{"comment":"The manuscript contains many spacing artifacts (e.g., 'T o', 'AL EA', 'g e ne rat e') that appear to originate from text extraction; a careful copyedit is needed before publication.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the journal's scope in computer science education or learning technologies. I see no grounds for rejection: the central claims are cautiously stated and the limitations section is unusually transparent. However, the missing rater-reliability information is load-bearing for the main counts, and the relational-annotation result lacks any quantitative measure. A revision that reports the rater panel, inter-rater agreement, and a simple annotation-correctness metric would allow me to endorse acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing here is the negative result: GPT-4 Turbo with retrieval-augmented generation can produce structurally annotated sTeX questions, but relational annotations fail and many questions need human filtering. That is worth knowing. It is also the first paper I have seen that combines RAG with sTeX semantic annotations for course-specific question generation, and the explicit split between structural and relational annotation success is a nice framing.\n\nThe paper is honest and well described. The prompt iteration is transparent, the limitations section is candid, and the example of a wrong question that reinforces the denying-the-antecedent fallacy is a genuinely instructive illustration of why expert review is needed. This is not a hype paper; it sets modest claims and mostly sticks to them.\n\nThe soft spot is exactly where the stress-test note points: the central counts — 28/30 fit, 27/30 solvable, 11/30 with content errors — come from an expert survey that never reports how many experts rated the questions, whether any question was rated more than once, or any inter-rater agreement statistic. The authors also concede there were no control items and no student validation. With a single rater per item, every count is a proxy for that one rater's severity; with multiple raters, unreported disagreement could hide real instability. The relational-annotation failure is described qualitatively, without per-question accuracy counts, so that part is even harder to verify. These gaps do not overturn the broad qualitative conclusion, but they do cap how strongly the numbers can be stated. The sample is small (30 questions, six topics, one model, one point in time), and no data or code are provided, so independent verification is not possible.\n\nThis paper is for researchers in automated question generation, adaptive learning, and computing education. It is a solid, honestly reported empirical study with a useful negative result, and it deserves a serious referee. I would send it to review, but with a request for the rater details and some acknowledgement of how the missing reliability figures constrain the quantitative claims. If I cite it, it will be for the qualitative finding, not for the specific fractions.","headline":"A useful, honest negative result about LLM question generation, but the expert evaluation needs more methodological rigor before the numbers can be trusted.","tokens_in":574,"tokens_out":1678,"would_cite":true,"duration_ms":32083,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4 can add structural markup to course questions, but not reliable concept links.","keywords":["automated question generation","large language models","retrieval-augmented generation","semantic annotation","sTeX","computer science education","adaptive learning","Bloom's taxonomy"],"falsifier":"Give the same 30 questions to a second panel of course experts, or to students who have completed the course, and compare solvability and content-error judgments; if agreement is low, the reported quality counts do not generalize.","tokens_in":15947,"feed_emoji":"🎓","tokens_out":3864,"duration_ms":36349,"temperature":0.7,"pith_summary":"This paper tests whether a large language model can generate quiz questions for a specific university course that are both pedagogically sound and annotated with the semantic information an adaptive learning system needs. Using retrieval-augmented generation, the authors feed GPT-4 Turbo relevant sections of a symbolic AI lecture and ask it to produce sTeX-annotated multiple-choice or fill-in-the-blank questions targeting Bloom's 'understand' level. The central result is mixed: structural annotations (question environments, correct-answer markers, objectives) are produced correctly in almost all cases, but relational annotations linking text to ontology symbols rarely work, and 11 of 30 expert-rated questions contain content errors. The authors conclude that LLMs can contribute to a pool of course-specific learning materials, but current performance requires significant human filtering and validation.","feed_headline":"GPT-4 writes course questions but fails at concept links","feed_subtitle":"Structural annotations succeed; relational links and question quality still need heavy human review.","key_machinery":"The load-bearing mechanism is the sTeX/OMDoc annotation framework paired with retrieval-augmented generation. sTeX marks up LaTeX documents with two kinds of annotations: structural annotations (the sproblem environment, mcb/scb choice blocks, fillinsol blanks, and objective declarations) that are schematic and independent of specific content, and relational annotations (symbol and module references that link text to concepts in an ontology). The pipeline injects the course's own sTeX source into the prompt, asks GPT-4 Turbo to produce annotated questions targeting the 'understand' level of Bloom's taxonomy, and then has domain experts rate the output on fit, solvability, clarity, relevance, feedback quality, and format alignment. The distinction between structural and relational annotations is what carries the central evaluation: the first works, the second does not.","core_discovery":"The paper's central claim is that a state-of-the-art LLM, GPT-4 Turbo, can generate course-specific quiz questions with correct structural sTeX semantic annotations, but fails at relational semantic annotations that connect question text to concepts in an ontology, and the generated questions frequently miss educational quality standards. In an expert evaluation of 30 questions across six topics from an 'Artificial Intelligence I' course, structural annotations (question environments, answer markers, objectives) were almost always correct, while relational annotations were rarely usable despite retrieval-augmented generation and function-calling attempts. Content errors appeared in 11 of 30 questions, often in answer options and feedback, including at least one question that actively reinforced a logical fallacy (denying the antecedent). The authors conclude that such LLM output can feed a pool of learning materials only under significant human review, and that the human-in-the-loop remains essential.","pith_inferences":[],"forward_implications":["Course-specific question generation is feasible with retrieval-augmented generation: experts rated 28 of 30 questions as fitting the teaching material and 27 of 30 as solvable from it.","Because relational annotation fails, any practical adaptive-learning pipeline must treat concept linking as a separate step from question writing rather than expecting the LLM to do both.","The generated question pool still needs expert review: 11 of 30 questions contained content errors, some of which actively reinforce common misconceptions.","LLMs strongly prefer multiple-choice and single-choice formats and rarely produce fill-in-the-blank questions when targeting the understand dimension.","Feedback is a persistent weak point: it is often missing or merely rephrases the wrong answer, limiting the learning value of otherwise usable questions.","The model's observed difficulty with 'apply'-level questions suggests that asking for deeper cognitive dimensions will require more than prompt tuning.","A hybrid design that uses the LLM only for question text and a deterministic ontology lookup for relational annotations may sidestep the main failure mode the paper identifies.","The expert-survey evidence would be strengthened by inter-rater agreement and student testing, both of which the paper notes are absent, and both are needed before the quality counts can be treated as stable."],"supporting_citations":[{"why":"GPT-4 Technical Report, the model used because of its large context window and measured performance.","marker":"[35]"},{"why":"sTeX3 system description, defines the markup that distinguishes structural from relational annotations.","marker":"[31]"},{"why":"OMDoc ontology, the underlying symbol/module framework that relational annotations must target.","marker":"[32]"},{"why":"Systematic review of automatic question generation, provides the background on methods and the gap concerning understanding-level questions.","marker":"[12]"},{"why":"Earlier LLM-based generation of programming exercises, a baseline this work extends toward course-specific semantic annotation.","marker":"[11]"},{"why":"Evaluation of GPT-4 and GPT-3 for MCQ generation, a comparison point for the newer model's question quality.","marker":"[20]"},{"why":"Review of AQG evaluation metrics, motivating the choice of expert-based evaluation.","marker":"[24]"},{"why":"Bloom's revised taxonomy, defines the cognitive dimension used when targeting 'understand'.","marker":"[21]"}],"fun_headline_variants":["GPT-4 writes quiz questions, but concept links fail","LLM question generation: structural wins, relational loses","AI-crafted exam questions need human polish for quality","Quiz questions from GPT-4: good structure, poor meaning","LLMs generate course questions, but human review is a must"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results stand on the assumption that the experts' ratings of the 30 questions are consistent and representative; the paper reports no inter-rater agreement, no student evaluation, and no control items, so the counts (28/30, 27/30, 11/30) may be specific to this course and this model.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4 writes quiz questions, but concept links fail","LLM question generation: structural wins, relational loses","AI-crafted exam questions need human polish for quality","Quiz questions from GPT-4: good structure, poor meaning","LLMs generate course questions, but human review is a must"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000721,"raw_usage":{"total_tokens":3212,"prompt_tokens":898,"completion_tokens":2314,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":2233}},"tokens_in":514,"tokens_out":2314,"duration_ms":15854,"temperature":1.0,"reasoning_tokens":2233,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:39:42.048036+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same 30 questions to a second panel of course experts, or to students who have completed the course, and compare solvability and content-error judgments; if agreement is low, the reported quality counts do not generalize.","supporting_citations":[{"cited_title":"System Description: s T eX3 – A LATEX-based Ecosystem for Semantic/Active Mathematical Docu- ments","cited_arxiv_id":null,"evidence_quote":"sTeX3 system description, defines the markup that distinguishes structural from relational annotations."},{"cited_title":"OMDoc – An open markup format for mathematical documents [Version 1.2]","cited_arxiv_id":null,"evidence_quote":"OMDoc ontology, the underlying symbol/module framework that relational annotations must target."},{"cited_title":"A Systematic Review of Automatic Question Generation for Educational Purposes","cited_arxiv_id":null,"evidence_quote":"Systematic review of automatic question generation, provides the background on methods and the gap concerning understanding-level questions."},{"cited_title":"Automatic Generation of Programming Exercises and Code Explanations Using Large Language Models","cited_arxiv_id":null,"evidence_quote":"Earlier LLM-based generation of programming exercises, a baseline this work extends toward course-specific semantic annotation."},{"cited_title":"Generating Multiple Choice Questions for Computing Courses Using Large Language Models","cited_arxiv_id":null,"evidence_quote":"Evaluation of GPT-4 and GPT-3 for MCQ generation, a comparison point for the newer model's question quality."},{"cited_title":"Automatic Question Generation: A Review of Methodologies, Datasets, Evaluation Metrics, and Applications","cited_arxiv_id":null,"evidence_quote":"Review of AQG evaluation metrics, motivating the choice of expert-based evaluation."},{"cited_title":"A T axonomy for Learning, T eaching, and Assessing: A Revision of Bloom’s T axon- omy of Educational Objectives","cited_arxiv_id":null,"evidence_quote":"Bloom's revised taxonomy, defines the cognitive dimension used when targeting 'understand'."}],"review_version":1}