{"id":"3075791e-c334-47e3-8667-d0901d2ac668","arxiv_id":"2608.04006","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A co-design study shows that making LLM trustworthiness metrics visible to learning engineers modestly increases agreement when choosing between AI tutor responses.","lead":"Learning engineers who build LLM tutors co-designed five metrics and visualizations that flag pedagogical problems in chatbot responses. In a 30 matchup evaluation, showing these metrics increased expert agreement on the best LLM response, though the increase was modest and below standard reliability.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Automated violation pipeline is unvalidated in educational contexts; if flags are inaccurate, the observed IRR increase may reflect anchoring on shared noise rather than trustworthiness.","rationale":"The reader's weakest assumption identifies the same load-bearing risk: the automated violation pipeline has not been validated for educational content, so the intervention may not actually make trustworthiness explicit. I agree with that assessment, and it is the most fundamental threat to the central claim. Even if the IRR increase were statistically robust (which is itself uncertain, as the reader notes — α0=0.3987 vs α1=0.4931 has no CI or test), it would not establish the claim if the displayed metrics are inaccurate. The paper is transparent about this limitation in §8.3 and about fitting thresholds in §5.3, which is a credit to the authors, but transparency does not remove the need for a validation check. A labeled validation set and human–pipeline agreement comparison would directly test whether the flags are accurate enough to serve as decision evidence. If the test shows high accuracy, the concern is retired; if not, the headline claim should be treated as a design hypothesis rather than a demonstrated effect. Because the reader already reached CONDITIONAL on essentially these grounds, my stress-test does not change the verdict.","tokens_in":45945,"tokens_out":6584,"duration_ms":65793,"concrete_test":"Run a validation substudy: take the 150 generated responses (or at least a stratified sample of 40–60 spanning all five prompts), recruit 3–5 learning engineers who did not design the pipeline, and have them independently label each of the 20 measures in Fig. 10 as violated/not violated using the paper's definitions. First report human–human agreement on a subset, then compute per-measure precision, recall, and F1 of the automated pipeline against the majority human label. If F1 is materially below 0.8 on high-usage measures (Misinformation, Hallucination, Feedback, Socratic, Scaffolding, Readability), or if pipeline–human agreement is no better than human–human agreement, the observed behavioral shifts cannot be attributed to trustworthy metrics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the metrics shown to raters actually measure pedagogical trustworthiness. The pipeline in §5.3 combines LLM claim extraction, DeBERTa textual entailment (labels accepted above 50% confidence), cosine similarity (threshold 0.5), and gpt-5 LLM-as-a-judge for the five Learning Methods, with thresholds for Participation and Diversity fit to the study dataset. §8.3 states explicitly that the measures' 'accuracy applied to educational contexts remains untested.' No precision/recall or agreement against expert human labels is reported for any of the 20 measures. If the flags are wrong, the 'With Metrics' condition is not displaying trustworthiness; it is displaying a possibly arbitrary overlay. The increased inter-rater agreement (α0=0.3987 vs α1=0.4931) and the shift toward Prompt 5/Prompt 2 could then be explained by raters anchoring on the same spurious badges, not by genuine calibration of trustworthiness. This is an internal-validity threat to the mechanism, not merely a missing benchmark.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a three-phase co-design study with learning engineers building an LLM-powered digital textbook. The authors first develop five trustworthiness metrics (Truthfulness, Diplomacy, Disposition, Learning Methods, Response Quality) comprising 20 measures, then design visualizations that map metric violations onto LLM responses, and finally evaluate the tools in an A/B prompt-evaluation tournament with 12 learning engineers rating 30 match-ups. The headline result is that making trustworthiness metrics visible increased inter-rater reliability (Krippendorff's alpha from 0.3987 without metrics to 0.4931 with metrics) and helped participants reconcile conflicting rubric objectives. The paper also contributes design guidelines for future LLM evaluation tools.","tokens_in":46207,"tokens_out":4261,"duration_ms":47180,"significance":"If the headline effect holds, this is a useful contribution to HCI and learning-engineering practice: it provides a concrete, domain-grounded operationalization of LLM trustworthiness for educational evaluation and a careful design-study template. The study has real strengths: it is ecologically valid (real student conversations, authentic prompt templates, within-subjects alternating interface conditions), it is longitudinal and co-design-based, and the authors are explicit about several key limitations—notably that the automated measures' accuracy in educational contexts is untested (§8.3) and that Participation and Diversity thresholds were fitted to the study dataset (§5.3). The qualitative findings—metrics acting as a procedural checklist, visualizations supporting trade-off reasoning—are credible and well supported by participant quotes. However, the central quantitative claim currently rests on an unvalidated measurement pipeline and an unsupported statistical comparison, which are the main barriers to accepting the paper's conclusions.","major_comments":[{"comment":"The central claim presupposes that the 20 measures produce valid violation flags that learning engineers can use as trustworthy evidence. Section 8.3 states that 'their accuracy applied to educational contexts remains untested,' and no precision/recall, expert-label agreement, or any other validation is reported for any measure. The observed IRR increase (α0=0.3987 to α1=0.4931) is therefore also consistent with raters anchoring on a shared but possibly arbitrary overlay rather than on genuine trustworthiness information. Please add a validation study—e.g., expert labels on a held-out sample of responses with per-measure agreement, or at minimum a sensitivity analysis showing that the IRR effect survives when flagged-but-questionable responses are removed.","section":"§5.3, §8.3"},{"comment":"The primary quantitative claim is the difference between α0=0.3987 and α1=0.4931. No confidence interval, bootstrap, or significance test is reported for either alpha or their difference, despite §4.3.6 adopting bootstrapped 95% CIs and the 'CIs do not overlap' rule. Both values are also below the conventional 0.67 reliability threshold, so the increase is from 'poor/moderate' to 'still poor/moderate.' Please report bootstrap CIs for each condition and a proper test of the difference (e.g., bootstrap percentile or permutation test); if the CIs overlap, the headline claim should be weakened accordingly.","section":"§7.1"},{"comment":"The Participation and Diversity violation thresholds were, in the authors' own words, 'set from the distribution of our own dataset rather than an external standard' and 'require recalibration before being applied to another corpus.' Because the same dataset is used in the tournament, the 'With Metrics' condition displays flags that are partly fitted to the evaluation set. This is a circularity: the increase in agreement is not yet evidence for trustworthiness metrics in general, only for this dataset-calibrated instantiation. Please report thresholds derived from prior work or via cross-validation, and test the robustness of the IRR difference to reasonable threshold variations.","section":"§5.3 (Fig. 10Q,R)"},{"comment":"Four of the twelve tournament raters participated in the earlier co-design phases for the metrics and visualizations (and, per §4.3.3, the additional eight were recruited later). Although participants had not seen the final visualizations, the co-designers helped define the design goals and measures; their prior investment may create demand characteristics or heightened fluency with the tool that the newly recruited raters lack. Please report whether the IRR increase (α0 vs. α1) holds when restricting the analysis to the eight non-co-design participants, or explicitly discuss this confound in the limitations.","section":"§4.3.3"}],"minor_comments":[{"comment":"The time-on-task comparison is reported descriptively (median 79.55s vs. 62.95s; 6/12 slower). This is fine as an exploratory result, but it would fit the paper's own §4.3.6 statistical framework to report bootstrap CIs for the medians or means.","section":"§7.2 / Fig. 16"},{"comment":"The 'Ratio of (higher−lower)-violation responses picked' values are shown with CIs in Fig. 17, but Fig. 19 is presented without CIs or raw counts. Adding counts would help readers assess the strength of the 'selected in spite of violation' patterns.","section":"Fig. 17 / Fig. 19"},{"comment":"The Dale-Chall readability formula and its constants are given only inside the figure. Since this is one of the 20 measures, please also define it in the text or a numbered equation so readers can reproduce the computation.","section":"Fig. 10O"},{"comment":"The CCS Concepts line still contains 'Do Not Use This Code' placeholders; this should be corrected before publication.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper's qualitative core is strong, and the quantitative shortcomings are fixable with additional analyses rather than fatal. I recommend major revision: the authors should validate or bound the violation pipeline, provide a proper statistical treatment of the IRR difference, and address the dataset-fitted thresholds and co-designer-rater overlap. If the validation cannot be done, the central claim should be reframed as a design study about perceived usefulness rather than a demonstrated increase in reliability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a genuine design study contribution, not a fraud. The five metrics / twenty measures co-designed with learning engineers, the overlapping-span visual encoding, and the deployment in a real prompt tournament are all real work. The empirical payoff is real but weaker than the abstract suggests. The IRR increase (0.3987 to 0.4931) is small, both values sit below the conventional 0.67 reliability threshold, and no significance test or confidence interval is reported. And the measure pipeline—claim extraction, DeBERTa entailment, gpt-5 judge, fitted thresholds—is explicitly admitted in Section 8.3 to have accuracy that \"remains untested\" in educational contexts. If the flags are wrong, raters are reacting to shared noise, so the stress-test concern lands; it is not manufactured. To the authors' credit, they disclose this limitation rather than hiding it, and the qualitative analysis and design guidelines stand on their own. The strongest parts are the workflow findings—metrics acting as a procedural checklist and visualizations helping people reconcile internal reasoning—and the honest reporting of low alphas. Soft spots: two thresholds (Participation, Diversity) are fit to the same dataset used for evaluation; the \"most trustworthy\" prompt (Prompt 5) was not the overall winner, so rubric preferences still dominated; and the sample is twelve raters, one task, one model, one institution. These are limits, not fatal flaws. The paper deserves a serious referee. Ask for a significance test or CI for the IRR difference, and an appendix reporting precision/recall of the violation flags on a held-out set, even a small one. That would turn a solid design study into a defensible empirical claim.","headline":"A well-executed co-design study with a plausible but under-supported central claim: the violation pipeline is unvalidated, so the observed IRR increase could be anchoring on shared noise.","tokens_in":46679,"tokens_out":1889,"would_cite":true,"duration_ms":20724,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that making trustworthiness visible—as five metrics and traceable visualizations—raises expert agreement when learning engineers evaluate LLM responses, and argues for hybrid human-in-the-loop evaluation.","keywords":["LLM evaluation","trustworthiness","education technology","co-design","visualization","prompt tournament","inter-rater reliability","learning engineers"],"falsifier":"Run the same tournament with the same interfaces but with metric flags computed on shuffled or randomly mislabeled responses. If expert agreement rises as much with meaningless flags, the effect comes from having any shared display, not from trustworthy signal. Also, compare the automated flags to expert human judgments on the same 30 responses: near-chance agreement would directly undermine the claim.","tokens_in":45852,"feed_emoji":"📊","tokens_out":6079,"duration_ms":62470,"temperature":0.7,"pith_summary":"This paper claims that making trustworthiness visible changes how learning engineers evaluate LLM responses in educational technology, and in the direction of better agreement. Through a year-long co-design with engineers building an LLM-powered digital textbook, the authors produced five trustworthiness metrics (20 measures) and visualizations that trace each violation to the sentence that caused it. In a prompt tournament with 12 learning engineers rating 30 real learner-sourced A/B match-ups, expert agreement rose when metrics were displayed, and engineers reported that the metrics acted as a checklist, surfacing pedagogical risks they had overlooked and making trade-offs between objectives explicit. The paper argues this supports hybrid evaluation: automated metrics screen for clear violations while humans resolve ambiguous, context-dependent cases. If right, trustworthiness is not just a benchmark property of a model but a workable decision-support lens for domain experts.","feed_headline":"Trustworthiness metrics lift expert agreement on LLMs","feed_subtitle":"In a 12-engineer prompt tournament, agreement rose from .40 to .49 once metrics and violation maps were shown.","key_machinery":"The load-bearing mechanism is a per-response violation pipeline built on claim extraction and textual entailment. Each response's claims are checked against the textbook passage, the learner's summary, the chat history, and the latest learner message; entailment labels of entails/contradicts/neutral determine whether a claim counts as misinformation, hallucination, sycophancy, and so on—with neutral treated as the signature of an unsupported claim. Learning-method measures are scored by an LLM-as-a-judge against explicit pedagogical principles, and response-quality measures come from validated conversational features. The visualizations then map each flagged span back into the response text,","core_discovery":"The central claim is that operationalizing LLM trustworthiness as a set of violation metrics plus text-traceable visualizations improves the reliability of expert evaluation of LLM responses in education. The authors co-constructed five metrics—Truthfulness, Diplomacy, Disposition, Learning Methods, and Response Quality—composed of 20 measures, each computed per response by an automated pipeline. They then ran an LLM prompt tournament in which 12 learning engineers chose between paired responses, half the time with the metrics and visualizations visible. Expert agreement was higher with metrics visible (alpha 0.4931 vs 0.3987), and the distribution of 'best' responses shifted toward lower-vi","pith_inferences":["The agreement gain may come less from new information than from a shared reference point: any consistent, inspectable flag set might produce part of the effect, and a condition with plain-text metric summaries would isolate what visualization adds.","The same co-designed metric set could plausibly be adapted to other high-stakes LLM domains where experts disagree on outputs, since the mechanism is general rather than education-specific.","Because the study used one model and one dialogue task, the durability of the effect under repeated use and across models is open; a longitudinal replication would clarify whether trust in the measures grows or decays.","The dataset-derived thresholds (e.g., Participation, Diversity) should be recalibrated per corpus; without that, transferred deployments risk spurious flags that could erode the agreement benefit."],"forward_implications":["If the claim holds, prompt-evaluation panels in education can raise agreement without imposing a single weighting of criteria.","Automated violation flags can serve as a first-pass screening layer, with engineers resolving ambiguous cases.","The metric set offers a reusable starting point for education-specific LLM evaluation that can be adjusted to other pedagogical contexts.","Comparison-oriented visualizations (summary overview, glyphs, donuts) are the features engineers actually use; single-metric encodings like opacity and underlining need refinement.","Evaluation tools that show trustworthiness alongside rubric criteria can change which prompt wins, since the rubric winner may be the least trustworthy."],"supporting_citations":[{"why":"Supplies the eight-dimension trustworthiness framework and benchmark suite that the paper adapts into education-specific metrics.","marker":"[30]"},{"why":"Provides the prompt-tournament method used as the evaluation vehicle in Phase 3.","marker":"[27]"},{"why":"Supplies the LLM-as-a-judge technique used to score the five learning-method measures.","marker":"[71]"},{"why":"Provides the fine-tuned textual-entailment model that labels claims as entailed, contradicted, or neutral.","marker":"[25]"},{"why":"Supplies the definition of a claim used by the extraction pipeline.","marker":"[23]"},{"why":"Supplies the conversational-feature library behind the Response Quality measures such as Certainty, Readability, and Participation.","marker":"[29]"},{"why":"Defines SERT, one of the five learning methods whose principles the LLM-as-a-judge checks.","marker":"[43]"},{"why":"Supplies the Dale-Chall readability formula used for the Readability measure.","marker":"[6]"},{"why":"Supplies the design-study methodology that structures the longitudinal co-design process.","marker":"[52]"}],"fun_headline_variants":["Violation maps boost expert accord on LLM responses","Trust metrics raise rater agreement in LLM evaluations","Co-designed dashboards improve LLM trust checks","Metrics plus visuals sharpen LLM response grading","For LLM reviewers, trust metrics lift consistency"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The whole finding rests on the assumption that the automated flags—built from claim extraction, textual entailment, an LLM judge, and fitted thresholds—are accurate enough to be trusted as evidence; the paper states their accuracy in educational contexts is not yet tested.","fun_headline_variants_meta":{"raw":{"variants":["Violation maps boost expert accord on LLM responses","Trust metrics raise rater agreement in LLM evaluations","Co-designed dashboards improve LLM trust checks","Metrics plus visuals sharpen LLM response grading","For LLM reviewers, trust metrics lift consistency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000467,"raw_usage":{"total_tokens":2145,"prompt_tokens":705,"completion_tokens":1440,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":1382}},"tokens_in":449,"tokens_out":1440,"duration_ms":10961,"temperature":1.0,"reasoning_tokens":1382,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T04:12:42.885203+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same tournament with the same interfaces but with metric flags computed on shuffled or randomly mislabeled responses. If expert agreement rises as much with meaningless flags, the effect comes from having any shared display, not from trustworthy signal. Also, compare the automated flags to expert human judgments on the same 30 responses: near-chance agreement would directly undermine the claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the eight-dimension trustworthiness framework and benchmark suite that the paper adapts into education-specific metrics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the conversational-feature library behind the Response Quality measures such as Certainty, Readability, and Participation."},{"cited_title":"McNamara","cited_arxiv_id":null,"evidence_quote":"Defines SERT, one of the five learning methods whose principles the LLM-as-a-judge checks."},{"cited_title":"Chall and Edgar Dale","cited_arxiv_id":null,"evidence_quote":"Supplies the Dale-Chall readability formula used for the Readability measure."}],"review_version":1}