{"id":"3595f505-e4a1-4f51-a03f-17910cd507f4","arxiv_id":"2411.16936","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fine-tuning Mistral-7B and Llama3-8B on a new 15,000-clue Italian dataset generated by GPT-4o makes them imitate GPT-4o's clue style much more closely, with a small human evaluation suggesting better quality.","lead":"The authors build a system that turns Italian Wikipedia articles into crossword clues using large language models, and they release a new Italian clue dataset. They then fine-tune two open models and report that the fine-tuned versions closely imitate the teacher model, GPT-4o, though the evidence for actual educational value is weaker.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of significant improvement rests on a 100-context human evaluation with ambiguous rater setup and no inter-annotator agreement or significance test; the non-circular evidence is weaker than the conclusion.","rationale":"The strongest claim is conditional on external validation. ROUGE versus GPT-4o has limited evidential value because GPT-4o generated the training targets; the paper itself acknowledges that ROUGE is not reliable for semantic quality. The human evaluation is therefore the load-bearing evidence. It is too underspecified: no clear rater count, no blinding, no inter-annotator agreement, and no significance testing. This is not an accusation of fabrication; the reporting simply does not yet demonstrate reliability. The appropriate verdict is CONDITIONAL, requiring additional human-evaluation evidence. I agree with the reader's identification of the core weakness; my formulation focuses on the specific missing statistics rather than the philosophical point about GPT-4o as the gold standard. The public dataset and released models are useful resources, so the contribution remains valuable; the conditional verdict should stand.","tokens_in":23,"tokens_out":2868,"duration_ms":87923,"concrete_test":"Re-run the human evaluation on the same 100 contexts with three independent native Italian-speaking raters, blind to model identity and to the paper's hypotheses, and report per-model rating distributions, Fleiss' kappa, and a paired test (e.g., Wilcoxon signed-rank) comparing each fine-tuned model with its base on clues from identical contexts. If kappa is below 0.4 or the fine-tuned vs. base comparison is not significant at p<0.05, the claimed improvement is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline result—that fine-tuning Mistral and Llama with Italian-Clue-Instruct significantly improves clue generation—must ultimately be established by the human evaluation, because the automatic metric (Table 1) compares models against GPT-4o-generated references, and GPT-4o also produced the training data. A higher ROUGE score against GPT-4o after fine-tuning mostly confirms that the models learned to imitate the teacher; it does not by itself establish better educational clues. The human evaluation in Section 4 is the only independent evidence, but it is too underspecified to carry the claim. It covers 100 contexts with 3 clues each, uses a five-level rating, and is reported only as aggregate counts in Figure 5. The text does not state the number of raters unambiguously ('a native Italian speaker, master student of linguistics, and PhD student in linguistics' can be read as one person or three), does not say whether raters were blind to model identity, reports no inter-annotator agreement, and provides no statistical test comparing base and fine-tuned distributions. With four models and only 300 clues, the visible differences in Figure 5 could easily be within sampling noise, especially for small counts in categories A and B. The conclusion's phrase 'significant improvements' is therefore not supported by the evidence as reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Italian-Clue-Instruct, a dataset of GPT-4o-generated Italian crossword clues derived from Wikipedia articles, organized into four clue types (unrestricted, bare noun phrases, definite determiner phrases, and copular sentences). The authors fine-tune Mistral-7B-Instruct-v0.3 and Llama3-8b-Instruct with LoRA on this dataset and evaluate the resulting models using ROUGE scores against GPT-4o references and a human rating study. The central claim is that fine-tuning yields 'significant improvements' in the models' ability to generate educational crossword clues from Italian texts. The paper also describes the dataset curation pipeline, prompt templates, and an example generated crossword.","tokens_in":11070,"tokens_out":4528,"duration_ms":37797,"significance":"If the central claim is substantiated, the Italian-Clue-Instruct dataset and the fine-tuned models would be a useful resource for Italian educational crossword generation, an area with few dedicated tools. The paper explicitly releases the dataset and models, which is a concrete contribution, and the linguistic grounding of the four syntactic clue types is a strength. However, the headline improvement claim currently rests on a circular automatic evaluation (ROUGE against GPT-4o, which also generated the training data) and an underspecified human evaluation, so the evidential value of the paper's experiments is limited until those issues are addressed.","major_comments":[{"comment":"The automatic evaluation is circular with respect to the central claim. The fine-tuned models were trained on GPT-4o-generated clues (Section 3, 'Generation of Educational Italian Clues') and then evaluated with ROUGE scores computed against GPT-4o-generated references on a test set of 200 contexts. Matching the teacher distribution is exactly what the fine-tuning procedure optimizes, so the reported ROUGE gains (e.g., Mistral-7B ROUGE-1 from 0.342 to 0.611) largely confirm imitation of GPT-4o rather than an improvement in clue quality. The paper should either explicitly reframe these scores as a fidelity-to-teacher measure or provide a non-circular evaluation (for example, human judgments or a metric with human-written references) before claiming 'significant improvements'.","section":"Section 4, Table 1"},{"comment":"The human evaluation is too underspecified to support the conclusion of significant improvements. It covers 100 contexts with 3 clues each across four models, but the text does not unambiguously state the number of raters ('a native Italian speaker, master student of linguistics, and PhD student in linguistics' can be read as one person or three), does not say whether raters were blind to model identity, reports no inter-annotator agreement, and provides no statistical test comparing the base and fine-tuned rating distributions. Given the small counts in categories A and B, the visible differences in Figure 5 could easily arise from sampling noise. The paper should report the full rating setup, agreement measures, and significance tests, and should avoid the phrase 'significant improvements' without such support.","section":"Section 4, 'Evaluation Results with the human evaluator' and Figure 5"},{"comment":"The dataset-quality evaluation is not robust enough to support the claim that the majority of GPT-4o-generated clues are high quality. The ROUGE-based assessment (average ROUGE-1, ROUGE-2, ROUGE-L of 0.159, 0.114, and 0.146) is acknowledged by the authors to be unreliable for semantic quality, and the human evaluation is described only as a randomly chosen subset of 100 articles with no inter-annotator agreement or detailed rater instructions. Figure 4 should be accompanied by the rating protocol, the number of raters and their agreement, and confidence intervals or a test of the distribution before the dataset's quality is asserted.","section":"Section 3, 'Evaluating quality of the Italian-Clue-Instruct Dataset'"},{"comment":"The methodology states that GPT-4o generation was performed 'with human validation for accuracy,' but the only human validation described later is the evaluation of 100 sampled articles in Section 3. If this 100-article subset is the intended validation, the paper should say so explicitly; if there was additional manual validation of the full dataset, its scale and procedure should be reported. As written, the claim of human validation is unsupported.","section":"Section 3, 'Italian-Clue-Instruct Data Collection Methodology'"}],"minor_comments":[{"comment":"The abstract says the dataset comprises 'over 30,000 entries' while Section 3 reports '15,000 unique clues'; these numbers should be reconciled.","section":"Abstract and Section 3"},{"comment":"The caption says 'using GPT-4 Turbo' while the text and prompts consistently refer to GPT-4o; the inconsistency should be corrected.","section":"Figure 1 caption"},{"comment":"There is a typo: 'prompt engeneering' should be 'prompt engineering'.","section":"Section 3, 'Formulation of Various Prompts'"},{"comment":"The sentence 'a new dataset called 1 has been compiled' contains a placeholder footnote reference where the dataset name should appear.","section":"Section 1, last paragraph"},{"comment":"The phrase 'Apart we aimed to elicit three specific types of clue' is missing a word and should read 'Apart from this, we aimed to elicit three specific types of clues'.","section":"Section 3, 'Formulation of Various Prompts'"},{"comment":"The phrase 'Token Distribution of unique Content' and the duplicated 'Number of tokens' axis label should be cleaned up for readability.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper's contribution is potentially useful for the Italian educational NLP community, and the release of the dataset and fine-tuned models is a concrete asset. However, the evaluation section does not yet support the headline claim; I would like the revision to strengthen the human evaluation and reframe or supplement the automatic metrics. I also note that the abstract overstates the results relative to the evidence presented. This fits the conference scope if revised accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is the Italian-Clue-Instruct dataset and the two LoRA-fine-tuned models, both released. That is a legitimate resource for anyone working on Italian educational NLP, and the three syntactic clue types (definite DP, bare NP, copular) are a thoughtful, linguistically grounded way to diversify clue surface forms. The prompts are carefully engineered and the data collection pipeline is described in enough detail to reproduce. On those grounds, the paper earns its place at a workshop like CLiC-it.\n\nThe problem is the evidence for the central claim. The automatic evaluation in Table 1 measures ROUGE against GPT-4o-generated references on a test set. Since GPT-4o also produced the training data, higher ROUGE after fine-tuning mostly confirms that the models learned to imitate the teacher. That is not independently evidence of better educational clues. The paper even acknowledges ROUGE is not reliable for semantic quality, but then uses it as the main quantitative result.\n\nThe human evaluation is the only non-circular evidence, and it is too thin to carry the weight. One hundred contexts, three clues each, a five-level rating, and Figure 5 shows only aggregate counts. The text says the evaluation was done by \"a native Italian speaker, master student of linguistics, and PhD student in linguistics,\" which I cannot parse as either one person or three. There is no inter-annotator agreement, no blinding to model identity, and no significance test. With four models and roughly 300 clues per model, the visible differences in categories A and B could easily be sampling noise. Calling the gains \"significant\" is not supported by the reported data.\n\nThere are also internal inconsistencies that a referee would want cleaned up: the abstract promises \"over 30,000 entries\" but the methodology says 15,000 unique clues; Figure 1 says GPT-4 Turbo while the text says GPT-4o; and the layout description in the introduction is garbled. None of these are fatal, but they signal carelessness.\n\nBottom line: the dataset and models are useful, and the idea is sound. The evaluation needs a real human study with multiple raters, agreement metrics, and statistical testing before the improvement claim is credible. I would not cite this in my own work unless I were specifically working on Italian educational crosswords, but I would send it to peer review because the resource is valuable and the flaws are fixable. A serious referee should ask for the human evaluation to be redone properly.","headline":"Useful Italian clue dataset and fine-tuned models, but the headline 'significant improvements' rests on a circular automatic metric and an underspecified human evaluation.","tokens_in":11593,"tokens_out":1864,"would_cite":false,"duration_ms":18708,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning LLMs on a new Italian clue dataset sharply improves crossword clue generation.","keywords":["Italian crossword generation","clue generation","large language models","fine-tuning","educational NLP","self-instruct","Italian-Clue-Instruct dataset","GPT-4o distillation"],"falsifier":"Have multiple independent raters blindly compare base and fine-tuned clues on a few hundred contexts under the paper's five-level scale and compute inter-annotator agreement; if the fine-tuned advantage shrinks or disappears, the improvement claim is not supported. Alternatively, run a controlled experiment where human solvers must identify the answer from each clue; if fine-tuned clues do not lead to higher solving accuracy than base-model clues, the educational benefit is not confirmed.","tokens_in":10594,"feed_emoji":"🧩","tokens_out":6748,"duration_ms":55813,"temperature":0.7,"pith_summary":"The paper claims that fine-tuning two open weights language models, Mistral-7B-Instruct-v0.3 and Llama3-8b-Instruct, on a newly built Italian dataset makes them markedly better at generating educational crossword clues from a given text. The dataset, called Italian-Clue-Instruct, was assembled from Italian Wikipedia introductions and contains over 15,000 clues written by GPT-4o in four syntactic styles: unrestricted, bare noun phrases, definite determiner phrases, and copular sentences. On a test set of 200 contexts, the fine-tuned models' ROUGE scores against GPT-4o references roughly doubled, and a human rater assigned mostly the top rating to the fine-tuned models' output. If these results hold, educators gain an open, free way to turn any Italian educational text into customized crossword puzzles, and the released dataset becomes a benchmark for Italian clue generation.","feed_headline":"Fine-tuned LLMs write better Italian crossword clues","feed_subtitle":"A 15,000-clue Italian dataset pulls Mistral and Llama close to GPT-4o-level clue quality.","key_machinery":"The load-bearing component is the Italian-Clue-Instruct dataset together with the prompt templates used to create it. Each dataset entry pairs an Italian Wikipedia context with a keyword, the future crossword answer, and clues written by GPT-4o. The prompts force a sequence of rewriting operations: resolve all pronouns, split the text into independent sentences, choose three sentences that best characterize the keyword, and rephrase those sentences as clues that never contain the keyword or any part of it. Four prompt templates enforce four structures: no format constraint, a bare noun phrase with no determiner, a definite determiner phrase headed by a definite article, and a copular sentence with elliptical subject (for example, 'è una salsa piccante tipica della Tunisia' for the answer 'Harissa'). The same data then fine-tunes the two small models via LoRA, with the checkpoint of minimum loss among the three training epochs selected for evaluation.","core_discovery":"The central discovery is that a purpose-built instruction dataset can transfer GPT-4o's clue-writing ability to much smaller, open-weights models. The authors built Italian-Clue-Instruct by taking the opening sections of Italian Wikipedia articles, filtering them, extracting keywords, and prompting GPT-4o with four distinct prompt templates that force the generated clue into a specified syntactic structure. They then fine-tuned Mistral-7B and Llama3-8b with LoRA on about 15,000 of these clues. After fine-tuning, Mistral's ROUGE-1 score against GPT-4o references rose from 0.342 to 0.611 and Llama's from 0.258 to 0.552, with similar jumps in ROUGE-2 and ROUGE-L. A human evaluation of 100 contexts with three clues per context found that the fine-tuned models received mostly 'A' ratings—coherent, valid clues matching context, answer, and structure—while the base models received many lower ratings. The paper presents this as evidence that the dataset is an effective teaching signal for educational Italian crossword generation.","pith_inferences":["The gains are measured as similarity to GPT-4o, so they show successful imitation rather than proven educational benefit; a blind comparison against human-written clues or a test of solver accuracy would evaluate the pedagogical claim directly.","The syntactic control suggests a testable extension: measure human solvers' accuracy and response time on the four clue types to see whether the predicted processing differences actually appear in crossword solving.","The human evaluation used a single rater on 100 contexts; replicating it with several raters and reporting inter-annotator agreement would establish whether the rating scale is reproducible."],"forward_implications":["Teachers and students can generate customized Italian crossword puzzles directly from course texts using models that run locally without API costs.","The four clue structures allow puzzles to vary in syntactic difficulty, connecting the generator to psycholinguistic findings about which structures are harder to process.","The public release of the dataset gives Italian computational linguistics a benchmark for educational clue generation, a resource that did not exist before.","The demonstrated recipe—distill a strong model's clues into a small instruction set and fine-tune open models on it—generalizes to other languages and puzzle types.","The fine-tuned models narrow the quality gap with proprietary GPT-4o on this specific task, making high-quality Italian clue generation more accessible."],"supporting_citations":[{"why":"Supplies the self-instruct framework that structures the automated generation of clues from a strong model.","marker":"[27]"},{"why":"The English-language Clue-Instruct predecessor that this work extends to Italian.","marker":"[21]"},{"why":"The authors' earlier Italian crossword generator with few-shot prompting, the baseline this fine-tuning approach improves upon.","marker":"[18]"},{"why":"LoRA, the parameter-efficient fine-tuning method used to adapt Mistral and Llama to clue generation.","marker":"[30]"},{"why":"Provides the crossword schema generation method used to turn generated clues into a printable puzzle.","marker":"[17]"}],"fun_headline_variants":["Open-weight LLMs approach GPT-4o in Italian crossword clues","Dataset lifts Mistral, Llama to near-GPT-4o clue quality","Fine-tuned small models write GPT-4o-quality Italian clues","Dataset transfers GPT-4o clue skill to open-source LLMs","Closed-to-open transfer builds better Italian crosswords"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation rests on the assumption that GPT-4o's clue style is the right gold standard: the models are trained on GPT-4o clues and then scored by how closely they match GPT-4o clues, so the reported improvements are improvements in imitation, not necessarily in educational usefulness.","fun_headline_variants_meta":{"raw":{"variants":["Open-weight LLMs approach GPT-4o in Italian crossword clues","Dataset lifts Mistral, Llama to near-GPT-4o clue quality","Fine-tuned small models write GPT-4o-quality Italian clues","Dataset transfers GPT-4o clue skill to open-source LLMs","Closed-to-open transfer builds better Italian crosswords"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1574,"prompt_tokens":994,"completion_tokens":580,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":490}},"tokens_in":610,"tokens_out":580,"duration_ms":5844,"temperature":1.0,"reasoning_tokens":490,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:42:58.532256+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have multiple independent raters blindly compare base and fine-tuned clues on a few hundred contexts under the paper's five-level scale and compute inter-annotator agreement; if the fine-tuned advantage shrinks or disappears, the improvement claim is not supported. Alternatively, run a controlled experiment where human solvers must identify the answer from each clue; if fine-tuned clues do not lead to higher solving accuracy than base-model clues, the educational benefit is not confirmed.","supporting_citations":[{"cited_title":"Zeinalipour, T","cited_arxiv_id":null,"evidence_quote":"Provides the crossword schema generation method used to turn generated clues into a printable puzzle."}],"review_version":1}