{"id":"04052a89-5c8e-49c1-bdff-9179b326de49","arxiv_id":"2508.06846","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Color-coding every phrase in an LLM response by factuality score beats a plain answer: users trust it more and find it easier to validate accuracy.","lead":"Researchers tested different ways to show which parts of an AI answer are true. People liked and trusted a design that color-codes every phrase by factuality, and found it easier to check the answer.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Subjective ease ratings may not reflect true validation accuracy; no objective measure reported.","rationale":"The reader's weakest assumption (reliability of factuality scores) is a valid external-validity concern, and I partially agree. However, my load-bearing concern is more internal: the reported outcome measures are subjective, so even with perfect factuality scores, the claim that the interface 'enhances scrutiny' is not supported without an objective performance measure. The reader's UNVERDICTED verdict remains appropriate because the abstract lacks the necessary detail; I identify this concrete gap and a test to fill it.","tokens_in":702,"tokens_out":2491,"duration_ms":25369,"concrete_test":"Inspect the paper's full method and data: does the experiment include an objective dependent variable, e.g., asking participants to flag which specific phrases are inaccurate, compared to ground truth? If not, run a replication (N≈50 per condition) measuring both subjective ease and objective detection accuracy against ground truth. If the all-phrase color-coded condition does not significantly outperform the no-style baseline on objective detection, the claim that it enhances users' ability to scrutinize outputs is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract reports that participants found the all-phrases color-coded style 'easier to validate accuracy' and that it 'enhances users' ability to scrutinize LLM outputs.' But the described experiments measured only ratings of trust, ease, and preference—no objective measure of whether users actually identified hallucinated phrases correctly. If the 'easier' rating does not correlate with improved detection performance, the central transparency claim is weakened: the design might only look helpful while not genuinely improving scrutiny. This is especially important because the factuality scores being visualized are the ground truth of the task; if those scores are noisy, biased, or miscalibrated (the abstract gives no information about how they were generated or validated), color-coding all phrases could actively mislead users into trusting false content. The paper's design guidelines for 'calibrating user trust' therefore depend on both (a) an objective validation of user performance, and (b) the reliability of the underlying factuality scores. Neither is established from the abstract.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two scenario-based experiments with 208 participants comparing interface styles for communicating per-phrase factuality scores in LLM responses. The central claim is that users prefer, trust, and find it easier to validate accuracy in a design where all phrases are color-coded by factuality score, compared to a baseline with no style applied. Based on these results, the authors offer design guidelines for calibrating user trust and improving scrutiny of LLM outputs.","tokens_in":915,"tokens_out":1934,"duration_ms":21172,"significance":"If the findings hold, the paper makes a useful empirical contribution to LLM transparency research by moving beyond factuality detection to the question of how to communicate score information to users. A user-centered evidence base for interface choices is valuable for practitioners. However, the significance is currently limited because the abstract reports only subjective ratings and does not establish that color-coding improves actual validation performance or that the underlying factuality scores are reliable enough to be visualized in this way.","major_comments":[{"comment":"The abstract concludes that the design 'enhances users' ability to scrutinize LLM outputs,' but the reported outcome measures are trust, ease in validating accuracy, and preference—all subjective ratings. No objective measure of whether participants actually identified hallucinated phrases more accurately is reported. Perceived ease may diverge from objective performance; the transparency claim needs at least one behavioral outcome (e.g., precision/recall in flagging incorrect phrases).","section":"Abstract (overall claims)"},{"comment":"The abstract does not state how the per-phrase factuality scores were produced or validated. The visualization's usefulness depends entirely on these scores being accurate, calibrated, and interpretable. If the scorer is noisy or biased, color-coding all phrases could mislead users into trusting false content. The authors must specify the scorer, its validation, and how its errors would be communicated or mitigated.","section":"Abstract (factuality score generation)"},{"comment":"The abstract says the results are 'statistically comparative' but provides no effect sizes, confidence intervals, or test statistics. From the abstract alone, the reader cannot assess the practical magnitude of the preference for the all-phrases color-coded condition. The full paper should report these details, and the abstract should include at least effect sizes or a summary confidence statement.","section":"Abstract (statistical reporting)"}],"minor_comments":[{"comment":"The baseline is described only as 'no style applied.' Please clarify whether this baseline showed no factuality information at all or a non-colored/plain textual format. This affects interpretation of the comparison.","section":"Abstract (baseline description)"}],"recommendation":"major_revision","confidential_remarks":"This is an abstract-only review; the full text may already address some points. The main risk I see is the gap between subjective ease/preference and objective scrutiny performance, which is central to the stated contribution. I would like the authors to add behavioral outcome measures and to state the provenance and validation of the factuality scores. These are fixable in a revision, so I am not recommending rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a worthwhile, well-scoped user study. The research gap—how to present factuality scores to users—is real, and the authors run two scenario-based experiments with 208 participants comparing design strategies. That alone puts it ahead of a lot of HCI papers that rally around a single anecdote. The finding that per-phrase color-coding beats baseline on trust, preference, and perceived ease is plausible and worth taking seriously.\n\nWhere I get cautious: the abstract reports no objective measure of whether users actually validated the responses more accurately. \"Easier to validate\" is a subjective rating, not a detection performance. A design can feel easy and still not improve scrutiny, or can even make users lazier. That’s a real threat to the transparency claim, not a pedantic point. If the paper ships with only self-reported ease, the central guidelines are over-reaching.\n\nThe other soft spot is the factuality scorer itself. The color-coding is only as good as the per-phrase scores being visualized. If those come from a noisy or miscalibrated model, then highlighting every phrase could actively mislead users into trusting false content. The abstract gives zero details on how the scores were generated, validated, or what accuracy they have. For a study whose whole point is to calibrate trust, the trustworthiness of the ground-truth signal matters.\n\nThat said, none of this is disqualifying. The concern about subjective ratings is the biggest one, and it is testable; a follow-up or revision with an objective verification task (e.g., asking participants to flag hallucinated phrases) would substantially strengthen the paper. I’d also want effect sizes and the actual factuality scores’ reliability reported.\n\nWho is this for: HCI researchers working on LLM transparency, UX designers building factuality indicators. It deserves a serious peer review, but with the expectation that the authors add objective performance data or at least clearly acknowledge the limitation. I would not desk-reject.","headline":"Useful design study on LLM factuality indicators, but the headline result rests on subjective ease ratings and an unvalidated underlying scorer.","tokens_in":1376,"tokens_out":1593,"would_cite":false,"duration_ms":15822,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Per-phrase color coding wins LLM users' trust","keywords":["LLM transparency","factuality scores","visual indicators","color-coding","user trust","hallucination communication","human-AI interaction","scenario-based experiment"],"falsifier":"Run the same two experiments but deliberately inject wrong factuality scores (e.g., random or inverted) for some phrases, then measure whether participants' trust and validation accuracy still favor the color-coded style, or whether the style produces misplaced confidence.","tokens_in":642,"feed_emoji":"🎨","tokens_out":2379,"duration_ms":23740,"temperature":0.7,"pith_summary":"This paper asks how to show users whether an LLM's answer is factual. The authors ran two scenario-based experiments with 208 participants comparing several visual styles for displaying factuality scores. They find that the design in which every phrase in the response is color-coded by its factuality score is preferred, trusted more, and rated easier to validate than a baseline with no styling. The result is a communication finding: when per-phrase scores exist, how they are shown determines whether users trust and scrutinize the output.","feed_headline":"Per-phrase color coding wins LLM users' trust","feed_subtitle":"Two experiments with 208 participants show color-coded phrases beat plain answers on trust and ease of checking.","key_machinery":"The central object is the 'all-phrases color-coded design,' a visual style that maps each phrase's factuality score to a color so the whole response becomes a factual heat-map. The mechanism is that this style gives a holistic, glanceable view of where a response is reliable and where it is not, compared to alternatives that only flag low-scoring spans or provide a single overall score.","core_discovery":"The paper claims that for LLM interfaces that have access to per-phrase factuality scores, the most effective way to communicate those scores to users is to color-code every phrase in the response according to its factuality score. In two scenario-based experiments with a total of 208 participants, this all-phrase color-coded design outperformed a baseline with no style on three measures: trust, preference, and perceived ease of validating response accuracy. The finding is about presentation, not detection — the authors take factuality scores as given and show that a full visual overlay, rather than highlighting only suspect phrases or showing a global score, best supports users' ability to","pith_inferences":["The paper's premise that factuality scores are trustworthy is untested; if the scorer is miscalibrated, color-coding could create false confidence in wrong answers.","Combining color with textual explanations or clickable evidence might further improve trust calibration beyond color alone.","The effect may depend on response length and density; too many colors could overwhelm users in longer outputs.","These findings could be tested against other annotation modalities (e.g., underlines, tooltips) and across different task types (e.g., factual question answering vs. creative writing)."],"forward_implications":["LLM interface designers should present phrase-level factuality scores as color overlays rather than only flagging low-scoring spans.","Users' trust calibration improves when every phrase carries a visible score, not just the 'bad' ones.","Ease of validating a response is higher with full phrase color-coding than with no styling, suggesting transparency features pay off.","The study provides concrete design guidelines for applications that already possess per-phrase factuality scores."],"supporting_citations":[],"fun_headline_variants":["Color-code every phrase to earn LLM trust","Full color overlay best for LLM factuality checks","All-phrase highlighting boosts LLM transparency","Color each phrase: users trust LLM more","Visual factuality coding wins user confidence"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The per-phrase factuality scores being visualized are accurate, and participants cannot see the underlying evidence, so the color-coding inherits the scorer's errors; if the scorer is noisy or biased, the color-coded design could mislead users more than a plain answer.","fun_headline_variants_meta":{"raw":{"variants":["Color-code every phrase to earn LLM trust","Full color overlay best for LLM factuality checks","All-phrase highlighting boosts LLM transparency","Color each phrase: users trust LLM more","Visual factuality coding wins user confidence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2191,"prompt_tokens":695,"completion_tokens":1496,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":1437}},"tokens_in":439,"tokens_out":1496,"duration_ms":10735,"temperature":1.0,"reasoning_tokens":1437,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:27:59.809510+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two experiments but deliberately inject wrong factuality scores (e.g., random or inverted) for some phrases, then measure whether participants' trust and validation accuracy still favor the color-coded style, or whether the style produces misplaced confidence.","supporting_citations":[],"review_version":1}