{"id":"f03c229c-a019-4b5d-85d0-34def1186ef8","arxiv_id":"2505.09923","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-dimensional, rubric-based LLM evaluation framework for question quality is proposed and illustrated on CAUS and SQUARE datasets, but its validation is preliminary.","lead":"This paper proposes a rubric-based system for judging whether questions are good, scoring them on appropriateness (fitting the context) and effectiveness (achieving the speaker's goal). It tests the system on three hand-made examples and two existing question datasets using an LLM as the judge.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Empirical support rests on an unvalidated LLM judge; without reported human-agreement statistics, the score patterns could reflect model bias rather than question quality.","rationale":"The reader's weakest assumption identifies exactly the load-bearing concern: the entire empirical validation hinges on an LLM judge whose agreement with human evaluators is asserted but never quantified. I agree that this is the central soft spot. The three-question validity test is illustrative at best: it shows the judge can rank obviously bad questions below an obviously good one under a rubric, which is not the same as showing the rubric measures question quality across real contexts. The CAUS/SQUARE demonstrations add descriptive patterns but no ground-truth comparison. Since the reader already assigned a CONDITIONAL verdict with this concern as the main condition, my stress-test does not move the verdict. The framework is conceptually coherent and the rubric is clearly described, so the appropriate outcome is not rejection but a requirement for human-agreement evidence, baselines, and statistical tests before the empirical claims can be accepted. My only slight extension of the reader's point is that the lack of human validation also weakens the model-selection claim (Methods, Evaluation Method) and makes the LLM-generated CAUS results potentially circular, but both are subsumed under the same missing measurement-validity evidence.","tokens_in":10485,"tokens_out":3225,"duration_ms":34317,"concrete_test":"Have at least three independent human raters score the same 150 CAUS and 150 SQUARE context-question pairs, plus the three validity-test follow-ups, using the paper's rubric with the same context variables. Compute human-human inter-rater reliability (e.g., Krippendorff's alpha) and LLM-human agreement (e.g., quadratic-weighted Cohen's kappa per sub-component). If the LLM's agreement with humans is not comparable to human-human agreement, or if human raters themselves disagree on appropriateness/effectiveness, the reported score patterns do not validate the framework. Reporting the omitted comparison of claude-3-5-sonnet versus the other three models on this labeled set would also directly settle the model-selection claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that the framework can 'assess both well-formed and problematic questions while adapting to varied contexts'—requires that the claude-3-5-sonnet judge's scores correspond to human judgments of question quality. The paper's only direct validity evidence is a three-question demonstration (Figures 1-2) showing that the judge assigns lower scores to two hand-crafted 'invalid' follow-ups than to the original. This does not establish validity: the items were constructed by the authors to be obviously deficient, so a prompted LLM could reproduce the intended ordering from the rubric descriptions alone; no human labels, agreement coefficients, or statistical tests are reported. The Methods section asserts that claude-3-5-sonnet was selected after 'comparative testing' with GPT-4, GPT-4-turbo, and Claude Opus because it showed 'the highest agreement with human evaluators,' but no agreement numbers or evaluation protocol are given. Without these, the CAUS and SQUARE score profiles (Tables 3-4) cannot be read as evidence that the framework measures question quality rather than the LLM's own stylistic preferences. Moreover, the CAUS questions were themselves LLM-generated; if the judge systematically favors LLM-like phrasing, the uniformly high clarity and respectfulness scores are partly an artifact of that preference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a rubric-based framework for evaluating question quality in human-AI interaction. It defines two main dimensions—appropriateness, understood as sociolinguistic fit to context, and effectiveness, understood as strategic goal achievement—and breaks each into three subcomponents: cohesion, answerability, and respectfulness for appropriateness; clarity, coherence, and informativeness for effectiveness. A five-point scoring rubric is provided for each subcomponent, with two dynamic contextual variables, ${answerer} and ${goal}, inserted into selected rubric descriptions. The authors report a small validity test using one legitimate and two deliberately invalid follow-up questions (Figures 1-2), then apply the rubric via claude-3-5-sonnet-20240620 to 150 questions from the CAUS dataset and 150 questions from the SQUARE dataset, reporting means and standard deviations in Tables 3-4. The paper concludes that the framework can assess both well-formed and problematic questions while adapting to varied contexts.","tokens_in":10703,"tokens_out":6108,"duration_ms":59848,"significance":"If the measurement instrument were properly validated, the framework would be a useful contribution: the conceptual distinction between appropriateness and effectiveness is clearly motivated by pragmatics and speech-act theory, the rubric is detailed and operationalized, and the contextual variables offer a sensible mechanism for semi-adaptive scoring. The authors also release code and detailed evaluation procedures in a public repository, which supports reproducibility. However, the current empirical evidence does not establish that the LLM judge's scores correspond to human judgments of question quality, so the central claim is not yet supported. At this stage, the main value of the paper is the rubric design and theoretical framing; the validation is too thin to support the abstract's claims of demonstrated assessment ability.","major_comments":[{"comment":"The manuscript states that claude-3-5-sonnet-20240620 \"was selected after comparative testing with GPT-4, GPT-4-turbo, and Claude Opus, showing the highest agreement with human evaluators,\" but no agreement coefficients, number of raters, annotation instructions, or comparison table are provided. Since all reported score patterns in Tables 3 and 4 are produced by this single LLM judge at temperature 0, the central claim that the framework \"assesses\" question quality cannot be separated from the claim that this model's scores match human judgments. Without human-agreement statistics, the results could equally reflect the model's stylistic preferences. Please report the full human evaluation protocol, inter-rater agreement (e.g., Cohen's kappa or ICC) for each rubric dimension, and the comparative agreement results that motivated the model choice.","section":"Methods, Evaluation Method"},{"comment":"The validity test in Figures 1-2 uses only three author-written questions, and the rubric descriptions already encode the intended verdicts. For example, Informativeness level 1 is defined as \"seeks irrelevant or speculative information,\" which is exactly the property assigned to FQ#1 and FQ#2, and Cohesion level 1 is defined as \"contextually misused cohesive markers.\" The observed ordering is therefore largely a restatement of the scoring rubric rather than an independent test of the metric. The paper's own statement that \"we validated only three questions\" is not an adequate validity argument. Please validate on a larger, independently annotated item set that includes non-extreme cases near the boundaries of the scale, with human judgments collected under a transparent protocol.","section":"Validity Test of the Evaluation Metric"},{"comment":"The empirical evaluation is statistically underpowered and faces a circularity risk. The CAUS dataset is authored by the same research group (Shin, Kim, & Ryu, 2024) and consists of LLM-generated questions, so if the judge favors LLM-like phrasing, the high clarity and respectfulness scores in Table 3 may reflect that preference rather than question quality. Tables 3 and 4 report only means and standard deviations, with no inferential tests comparing the first/third/fifth question sets or the three SQUARE categories, no comparison against existing QG metrics or human baselines, and no inter-rater reliability. Consequently, conclusions such as \"the ethical question set needs improvement\" (Results, Applying to Irrelevant and Ineffective Questions) are not statistically supported. Please add pre-defined hypotheses, appropriate significance tests or effect sizes, and at least one baseline metric.","section":"Methods and Results, CAUS/SQUARE datasets"},{"comment":"The abstract's claim that the framework can adapt \"to varied contexts\" is not supported by the evidence. The only context manipulation is the two goal values in the three-question validity test (Figure 2); in the CAUS and SQUARE applications the answerer and goal variables are fixed within each dataset. No evaluation shows how scores behave across a range of answerer/goal values, and no reliability analysis of the semi-adaptive criteria is presented. Please either provide systematic context-variation experiments or soften the claim to match the actual scope of the demonstration.","section":"Abstract and Discussion"}],"minor_comments":[{"comment":"The abstract says the framework can \"access both well-formed and problematic questions\"; this should be \"assess\".","section":"Abstract"},{"comment":"The definition of effectiveness contains the typo \"sucessfully\"; it should be \"successfully.\"","section":"Table 1"},{"comment":"The dataset name is given as \"SQAURE\" in the Methods and Results but as \"SQuARe\" in the reference list and \"SQUARE\" elsewhere; please standardize the spelling.","section":"Methods and Results"},{"comment":"\"maximum token of 1500\" should be \"maximum tokens of 1500,\" and the sampling procedure for the 150-question subsets (e.g., random seed, inclusion criteria) should be reported.","section":"Methods, Evaluation Method"},{"comment":"The footnote numbering is inconsistent: Footnote 1 for the speech-act definition appears in the Introduction, but footnotes 3-5 are mentioned in the Methods without corresponding numbered markers in the text; please align the footnote markers.","section":"Footnotes"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is framed as an empirical validation, but the empirical content is closer to an illustrative pilot. If the journal's bar for this format is a validated measurement instrument, the authors should be asked to substantially strengthen the human-agreement evidence and statistical analysis; otherwise, the paper might be better reframed explicitly as a rubric proposal with illustrative scores. The same-group CAUS dataset and the unreported model-selection comparison are the main concerns for independence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The conceptual core of this paper is genuinely useful. Defining question quality along appropriateness and effectiveness, with six concrete sub-components and dynamic context placeholders, gives question-generation and human-AI interaction researchers a structured way to think about evaluation that n-gram overlap cannot touch. The rubric in Table 2 is detailed enough to adapt to new contexts, and the idea of parameterizing answerer and goal is a legitimate extension of LLM-as-a-judge with rubric-based scoring. The literature review is fair and well grounded; the distinction between formal and functional competence is handled sensibly. Credit where it is due: the framework is clearly specified and the paper is honestly written, even flagging that only three questions were used in the validity test.\n\nThe soft spots are exactly where the stress-test note lands. The only direct validity evidence is three handcrafted questions, and those questions are constructed to be deficient by design. That does not demonstrate the framework measures question quality; it shows the judge can read the rubric and assign low scores to obviously bad examples. The Methods section asserts that claude-3-5-sonnet was selected over GPT-4, GPT-4-turbo, and Claude Opus because it had the highest agreement with human evaluators, but no agreement numbers, no protocol, and no human labels are reported. Without that, the CAUS and SQUARE score profiles in Tables 3 and 4 could reflect the LLM judge's stylistic preferences rather than any underlying quality. The CAUS dataset being authored by the same group further weakens the independence of that demonstration. I also note the paper's own statement that \"we validated only three questions\" is an honest concession, but it does not change the fact that the central empirical claim outruns the evidence.\n\nThat said, the conceptual contribution does not collapse. The rubric itself could be used by human raters, and the two-dimensional decomposition is a reasonable proposal. The problem is the paper claims to have demonstrated an automated evaluation system, and that demonstration is not yet supported. The fix is straightforward: report human agreement statistics (e.g., Cohen's kappa or ICC), compare against existing QG metrics, and test the context variables more systematically.\n\nThis paper deserves a serious referee. It is coherent, novel in its specific combination, and potentially useful to the QG and LLM evaluation subfield. But it should come back with a major revision requirement, not just minor tweaks. The empirical validation needs to be rebuilt from the ground up.","headline":"A clearly argued rubric framework for question quality, but the empirical support rests on a three-question validity test and an unverified LLM judge.","tokens_in":11214,"tokens_out":1546,"would_cite":false,"duration_ms":16512,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a good question is both appropriate and effective, and that a six-criterion rubric with two context variables lets an LLM judge score question quality automatically.","keywords":["question quality evaluation","large language model","LLM-as-judge","rubric scoring","appropriateness","effectiveness","question generation","pragmatics"],"falsifier":"Take the same 300 questions from the CAUS and SQUARE datasets, have a panel of human judges apply the paper's rubric to them, and compare the human scores with the LLM's scores using a weighted agreement measure. If the LLM's scores diverge systematically from human ratings—or if a question that human judges unanimously consider excellent scores low on all six criteria—the framework's claim to measure question quality collapses.","tokens_in":10297,"feed_emoji":"❓","tokens_out":5398,"duration_ms":47523,"temperature":0.7,"pith_summary":"Questioning is everywhere in human discourse and increasingly central to AI, yet until now almost no one has tried to define what makes a question good rather than merely encouraging people to ask more. The paper proposes that a good question is one that is both appropriate and effective: appropriate means it fits the conversational context and social norms, effective means it achieves the asker's intended goal. To operationalize this, the authors build a five-point rubric with six sub-criteria (cohesion, answerability, respectfulness; clarity, coherence, informativeness) and inject two dynamic context variables—who the answerer is and what the discourse goal is—into the rubric descriptions. They then use claude-3-5-sonnet as a judge at temperature 0 to score follow-up questions from two datasets, one of well-formed uncertainty-resolving questions and one of intentionally problematic sensitive questions. The score profiles separate the two groups and shift when the goal variable changes, which the authors read as evidence that the framework can assess question quality in a context-sensitive way.","feed_headline":"Question quality: two dimensions, six criteria, one AI judge","feed_subtitle":"A six-criterion rubric lets an LLM judge separate good questions from bad ones, adapting to who answers and what the asker wants.","key_machinery":"The load-bearing mechanism is a five-point analytic rubric with six sub-components, each anchored by one of five ascending descriptions from 'complete deficiency' to 'full achievement'. Two dynamic placeholder variables are embedded in the rubric text: ${answerer} (who will respond), inserted into the Answerability criterion, and ${goal} (the discourse purpose), inserted into Clarity and Informativeness. The rubric is delivered as a prompt to the claude-3-5-sonnet-20240620 model at temperature 0, which scores each follow-up question on all six criteria; the placeholders make the evaluation semi-adaptive, so the same sentence can be judged differently depending on who is answering and what the conversation is trying to accomplish.","core_discovery":"The paper's central claim is that question quality reduces to two measurable dimensions: appropriateness (sociolinguistic competence in context) and effectiveness (strategic competence in goal achievement), with 'a good question is both appropriate and effective' as the summary definition. Each dimension decomposes into three rubric sub-components, and the rubric is made context-dependent through two placeholder variables, ${answerer} and ${goal}, that modify the scoring descriptions for answerability, clarity, and informativeness. The authors claim that this rubric, used as a prompt for an LLM judge, produces scores that distinguish a legitimate follow-up question from deliberately misleading or off-topic versions, and that the same question receives different scores under different goals, demonstrating adaptability. They also claim that applying the rubric to two contrasting datasets reveals distinct, interpretable quality profiles—high clarity and respectfulness for well-formed questions, and specific deficits (e.g., low respectfulness and cohesion) for ethically problematic questions.","pith_inferences":["If the LLM judge's agreement with humans is as high as the paper implies, the framework offers a scalable proxy for human evaluation of open-ended dialogue questions, which the paper does not directly demonstrate with agreement evidence.","The two-dimensional structure suggests a testable typology: questions can be appropriate-but-ineffective (polite and on-topic but missing the goal) or effective-but-inappropriate (goal-achieving but rude or norm-violating); the paper's FQ#1 illustrates the former, and constructing a rude-but-informative example would stress-test the latter.","The framework's reliance on explicit goal and answerer variables implies that practical deployment requires a separate step—inferring the user's goal and the answerer's identity—before scores are meaningful, a step the paper treats as given."],"forward_implications":["Question generation systems can be selected or fine-tuned against rubric scores, giving them an explicit quality signal instead of similarity to reference questions.","The same six-criterion rubric transfers across domains by simply re-instantiating the ${answerer} and ${goal} variables.","Score profiles become explainable diagnostics: for example, low respectfulness combined with low cohesion identifies ethically problematic questions.","AI systems can be evaluated on their questioning ability separately from their answering ability, supporting a shift toward question-centric interaction design."],"supporting_citations":[{"why":"Supplies the formal vs functional linguistic competence distinction that motivates the authors' focus on pragmatics rather than grammar.","marker":"(Mahowald et al., 2024)"},{"why":"Source of the appropriateness and effectiveness competence dimensions that structure the entire evaluation framework.","marker":"(Spitzberg, Canary, & Cupach, 1994)"},{"why":"Provides the rubric methodology for constructing clear, consistent scoring criteria.","marker":"(Brookhart, 2018)"},{"why":"Establishes the LLM-as-a-judge paradigm that the paper adapts by embedding the rubric in a prompt.","marker":"(Li et al., 2024)"},{"why":"A recent rubric-based LLM evaluation study whose five-point scoring practice the authors align with.","marker":"(Ye et al., 2024)"},{"why":"The CAUS dataset used for validating the rubric on well-formed uncertainty-resolving questions.","marker":"(Shin, Kim, & Ryu, 2024)"},{"why":"The SQUARE dataset used for validating the rubric on problematic sensitive questions spanning three categories.","marker":"(Lee et al., 2023)"},{"why":"Provides established practices for five-point rubric development followed in the scoring design.","marker":"(Popham, 1997)"}],"fun_headline_variants":["Two metrics decide if a question is good: appropriateness and effectiveness","AI judge uses six criteria to score your question's fit and goal","Two dimensions, six criteria: the anatomy of a good question","Good question? Check appropriateness and effectiveness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire validation depends on the assumption that claude-3-5-sonnet, prompted with the hand-written rubric at temperature 0, gives scores that match human judgments of question quality; the paper reports it was selected for highest agreement with human evaluators but does not report any agreement numbers or the human-evaluation protocol.","fun_headline_variants_meta":{"raw":{"variants":["Two metrics decide if a question is good: appropriateness and effectiveness","AI judge uses six criteria to score your question's fit and goal","Two dimensions, six criteria: the anatomy of a good question","Good question? Check appropriateness and effectiveness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000792,"raw_usage":{"total_tokens":3457,"prompt_tokens":878,"completion_tokens":2579,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":2512}},"tokens_in":494,"tokens_out":2579,"duration_ms":18221,"temperature":1.0,"reasoning_tokens":2512,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:20:41.407379+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same 300 questions from the CAUS and SQUARE datasets, have a panel of human judges apply the paper's rubric to them, and compare the human scores with the LLM's scores using a weighted agreement measure. If the LLM's scores diverge systematically from human ratings—or if a question that human judges unanimously consider excellent scores low on all six criteria—the framework's claim to measure question quality collapses.","supporting_citations":[{"cited_title":", Ivanova, A A","cited_arxiv_id":null,"evidence_quote":"Supplies the formal vs functional linguistic competence distinction that motivates the authors' focus on pragmatics rather than grammar."},{"cited_title":", Canary, D J","cited_arxiv_id":null,"evidence_quote":"Source of the appropriateness and effectiveness competence dimensions that structure the entire evaluation framework."},{"cited_title":", Kim, D","cited_arxiv_id":null,"evidence_quote":"The CAUS dataset used for validating the rubric on well-formed uncertainty-resolving questions."},{"cited_title":"APACrefauthors \\ 1997","cited_arxiv_id":null,"evidence_quote":"Provides established practices for five-point rubric development followed in the scoring design."}],"review_version":1}