{"id":"41272923-35d9-4abd-9324-7b90851aaeb1","arxiv_id":"2505.21512","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Visualizations designed to increase transparency in an LLM-based knowledge graph query tool actually led users, including experts, to overtrust incorrect outputs.","lead":"Researchers built LinkQ, a tool that lets people ask questions of knowledge graphs in plain language while showing them the generated query and its steps. In a study with 14 practitioners, users trusted the tool's answers even when they were wrong, partly because the helpful visuals made the system seem reliable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal claim that visualizations caused overtrust is unsupported by the design: there is no no-visualization control, and the quoted overtrust rationalizations in §6.2 do not mention visual features. A controlled comparison is needed.","rationale":"The strongest and most novel claim is causal: visual transparency designed to aid verification instead increased overtrust. This claim is load-bearing because the paper's contribution to visualization research depends on the role of the visual components, not merely on observing overtrust in an LLM-assisted tool, which prior work already documents. The missing comparison is the weakest link: alternative explanations—general LLM authority, task difficulty, users' pre-existing dispositions toward LLMs, and demand characteristics from being told the system \"will not always get the correct answer\"—are not controlled. The reported quotes strengthen the existence of overtrust but not its visual cause; several key rationalizations are not tied to any visual element. This does not mean the finding is wrong; it means the current evidence supports a conditional conclusion. The quantitative accuracy benchmark in Section 4.4 supports the system's utility but is orthogonal to the trust attribution, and the small purposive sample with 3/14 co-developer participants further limits confidence. The reader's conditional verdict is appropriate; the missing control condition and artifacts should be supplied before the causal framing is accepted as established. No verdict change from the reader's CONDITIONAL is warranted, hence UNCHANGED.","tokens_in":22060,"tokens_out":4722,"duration_ms":45123,"concrete_test":"Run a pre-registered between-subjects replication of the Section 5 protocol using the same intentional-failure questions: (A) the full LinkQ interface, and (B) LinkQ with all visual mechanisms disabled, leaving the chat panel and results table only, while holding the underlying query pipeline fixed. Measure the proportion of incorrect outputs accepted as correct, self-reported confidence, and number of verification actions such as source clicks and query inspections. If acceptance and confidence are not higher in condition A than in condition B, the visualization-caused-overtrust claim is refuted; if they are higher, the causal attribution gains direct support. Even a within-subjects follow-up that asks participants to rate trust with each visual component toggled on and off would provide a weaker but informative check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract states that users \"tended to overtrust LinkQ's outputs due to its 'helpful' visualizations,\" but Section 6.2 offers only the hedged inference \"we believe that the added visual transparency within LinkQ's natural language interface may have resulted in overconfidence in the LLM's outputs.\" The study design cannot separate visualization-driven overtrust from baseline LLM authority: every participant saw the full visual interface, and all participants were told in advance that LinkQ could be wrong and were shown a live failure example. The affirmative evidence for a visual mechanism is thin. Section 5.2 quotes praise for the state diagram (\"shed light,\" \"more engaging\") and the entity-relation table (\"gives me more confidence in the query\"), but these comments concern comprehension and perceived transparency, not acceptance of incorrect outputs. The two concrete incorrect-answer episodes in Section 6.2—the Gladiator rationalization and the Google Chrome \"no vulnerabilities\" rationalization—contain no reference to any visualization; they are plausibility-based rationalizations that would likely occur with any LLM assistant. Section 5.4 reports that KG experts believed answers were correct when the query structure \"looked good,\" but this is an inference from query plausibility, not a demonstration that the visualization itself caused the belief. Without a condition that removes or alters the visualizations, the central contribution remains an interpretation rather than an empirically supported causal finding. The involvement of 3/14 participants in system development and the absence of shared artifacts are secondary but reinforce the need for a cleaner test.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents LinkQ, an LLM-assisted visual knowledge-graph exploration system with five visualization components: an LLM-KG state diagram, a query editor with explanations, an entity-relation table, a query structure graph, and a results graph. The authors report a quantitative benchmark on the Mintaka dataset showing that LinkQ's chained-prompting protocol achieves higher query-generation accuracy than a GPT-4 baseline, and a think-aloud qualitative study with 14 KG/LLM practitioners. The study's main findings are that users, including KG experts, tended to overtrust LinkQ's outputs and sometimes rationalized incorrect answers; that users with different KG and LLM expertise adopted distinct verification and workflow strategies; and that the results graph was unused in favor of the tabular results view. Based on these observations, the paper offers design implications and preliminary guidelines for LLM-based KG systems.","tokens_in":22292,"tokens_out":5963,"duration_ms":50336,"significance":"If the causal link between visual transparency and overtrust were established, the paper would make an important contribution to visualization and explainable-AI research, challenging the common assumption that transparency mechanisms uniformly increase critical scrutiny. The descriptive findings—that workflow strategies vary systematically with KG/LLM expertise and that the most popular visualization (the state diagram) fosters trust—are timely and relevant for designers of LLM-assisted visual analytics systems. The system itself is open-source, and the quantitative benchmark, while limited, supports the claim that the chained-prompting protocol improves query accuracy over a plain LLM. The qualitative data are presented with direct participant quotes, which is a strength and provides useful grounding for future work. However, the central causal claim that visualizations caused overtrust is not supported by the current study design, so the paper's main contribution is better characterized as a hypothesis-generating qualitative study than as an established causal finding.","major_comments":[{"comment":"The central claim that LinkQ's visualizations caused user overtrust is not supported by the study design. All participants saw the full visualization suite, with no condition that omitted or altered the visualizations, so the observed overtrust cannot be separated from the general authority of an LLM or from participants' prior dispositions. The two concrete incorrect-answer rationalizations in §6.2—the Gladiator and Google Chrome examples—make no reference to any visualization; they are plausibility-based rationalizations that would plausibly occur with a text-only LLM assistant. The quotes in §5.2 praising the state diagram and entity-relation table concern comprehension and perceived transparency, not acceptance of incorrect outputs, and the paper does not report how often overtrust occurred ('often' is not quantified). The abstract's claim that users 'overtrust LinkQ's outputs due to its helpful visualizations' should be softened to a hypothesis, or a controlled comparison (e.g., an LLM chat-only baseline) should be added to substantiate the causal role.","section":"Abstract; §5.1; §5.2; §6.2"},{"comment":"Three of the fourteen participants were involved in the initial stages of developing LinkQ. Their prior involvement creates a risk of demand characteristics and pro-system bias, yet the paper does not report whether their data were analyzed separately or whether any themes, including overtrust and visualization preferences, were robust to their exclusion. The authors should disclose this treatment or justify why these participants do not bias the reported findings.","section":"§5.1 (Participants)"},{"comment":"The selection of 'incorrect' questions for the qualitative study depends entirely on the quantitative evaluation reported in Table 1, but that table reports only aggregate percentages with no per-question results, confidence intervals, or significance tests. Moreover, the qualitative section does not specify which targeted questions each participant received or whether the pre-determined 'incorrect' label was verified in the actual session (e.g., by comparing with a ground-truth answer from the KG). Without this, the overtrust episodes cannot be interpreted as responses to genuinely incorrect outputs. Please provide the question set and the error-verification procedure.","section":"§4.4; §5.1 (Targeted questions)"}],"minor_comments":[{"comment":"Typo: 'aethestetically-pleasing' should be 'aesthetically pleasing.'","section":"§2.2"},{"comment":"The Gladiator example says 'a second one came out recently, so maybe that's when it won?'; the film Gladiator II was released in 2024, not 2023, so the participant's rationalization is even more implausible than the text implies. Please verify the details of the example.","section":"§6.2"},{"comment":"The statement that the graph visualization 'was never used' is strong; please clarify whether it was displayed by default or required an explicit user action, and whether 'not used' means not selected as the primary view rather than not seen.","section":"§5.2"},{"comment":"The paper would benefit from a short appendix listing the targeted questions used in the qualitative study, along with their ground-truth answers and the authors' classification (correct vs. incorrect). This would aid reproducibility and the interpretation of the overtrust episodes.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of a visualization venue and is likely to generate discussion. The main revision issue is the mismatch between the abstract's causal language and the hedged inference in §6.2. I recommend the authors be encouraged to reframe the contribution as a hypothesis for future controlled studies, or to add a minimal baseline condition. The self-citations (refs 41-43, 46) are concentrated in the authors' previous work on this system and do not bother me, but the paper's novelty should be clearly positioned relative to the prior LinkQ VIS 2024 paper (ref [41])."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: useful system paper plus an interesting, mostly well-hedged qualitative finding about overtrust. The abstract overclaims that visualizations caused the overtrust; the study design can't support that causal claim. That needs fixing, but the paper deserves refereeing.\n\nLinkQ itself is a real contribution: five purpose-built visual mechanisms for inspecting LLM-generated SPARQL queries, a chained-prompting protocol, and a small Mintaka benchmark where LinkQ beats plain GPT-4 by a wide margin. The qualitative study is well structured: think-aloud protocol, two KGs, targeted plus open-ended tasks. The expertise-dependent workflow findings are genuinely useful, and I buy the observation that people, including KG experts, rationalized wrong answers and that the tool's transparency created a sense of seeing the LLM work. That observation is worth publishing.\n\nSoft spots, in order. First, the central causal claim. The abstract says users overtrusted LinkQ's outputs \"due to its 'helpful' visualizations,\" but Section 6.2 only offers the belief that visual transparency \"may have resulted\" in overconfidence. There is no no-visualization control, no altered-visualization condition, and participants were primed that LinkQ could be wrong. The two concrete false-answer rationalizations in Section 6.2—Gladiator and Google Chrome—are plausibility-based and make no reference to any visual feature. The KG-expert \"query looked good\" remark in Section 5.4 does gesture at the query structure graph, so the visualization is not completely detached from the evidence, but the episode-level support is thin. This is an interpretation, not a demonstrated effect. Second, sample and independence: 14 participants, all AI/ML/cyber practitioners, 3 of them involved in building LinkQ, no shared artifacts in the version I read. Fine for a design study, limiting for a trust claim. Third, the quantitative results are reported without uncertainty or enough methodological detail in the main text. Fourth, the code/repo is redacted, so the open-source claim is uncheckable.\n\nCitation pattern is fine; the overreliance-in-XAI literature is cited, and the paper honestly positions itself relative to it. This is a paper for VIS/HCI and LLM-assisted analytics researchers.\n\nRecommendation: send to peer review. Ask the authors to align the abstract with the evidence, provide the artifacts, and ideally run a follow-up controlled comparison with the visualizations removed or altered. With those changes, this becomes a much stronger paper.","headline":"Solid design-study paper; the overtrust observation is real but the causal role of visualizations is asserted more strongly in the abstract than the design can support.","tokens_in":22857,"tokens_out":3311,"would_cite":true,"duration_ms":32065,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Visual transparency made users trust wrong LLM answers.","keywords":["knowledge graphs","large language models","natural language interfaces","SPARQL query generation","user trust","overtrust","visualization design","qualitative user study"],"falsifier":"Present matched groups of participants with the same deliberately wrong LinkQ answers, one group using the full visual interface and one using a chat-only version with identical LLM text but no state diagram, query graph, or entity-relation table; if wrong-answer acceptance rates are equal, the visualizations are not the cause of overtrust.","tokens_in":21847,"feed_emoji":"🤖","tokens_out":9217,"duration_ms":73204,"temperature":0.7,"pith_summary":"Knowledge graphs are powerful but hard to query, and LLM interfaces promise to lower that barrier. This paper asks whether the transparency visualizations added to such an interface help users judge the LLM's work or simply make them trust it more. The authors built LinkQ, a system that turns natural-language questions into SPARQL queries over knowledge graphs, and wrapped it in five visual mechanisms meant to expose where the LLM could go wrong. In a think-aloud study with 14 practitioners, users—including knowledge-graph experts—often accepted LinkQ's incorrect answers and cited the visualizations as evidence that the system was working. The paper argues that 'helpful' visual transparency can reduce critical scrutiny, and that these systems cannot be one-size-fits-all because workflows and trust vary with expertise and prior skepticism.","feed_headline":"Visual transparency made users trust wrong LLM answers","feed_subtitle":"In a 14-user study of LinkQ, 'helpful' visuals made even KG experts overtrust false LLM outputs.","key_machinery":"The load-bearing mechanism is LinkQ's prompting protocol together with its five visualization components: the LLM-KG state diagram, which shows which stage of the pipeline the system is in; the query editor, which pairs generated SPARQL with an LLM explanation; the entity-relation ID table, which gives human-readable labels for the entities and relations the LLM found; the query structure graph, which draws the graph pattern the query will traverse; and the results graph visualization, which embeds result tables in the query graph. These components are designed to surface the points where an LLM could hallucinate or misidentify data. The study's finding is that the same mechanisms supply the visual plausibility that lets users trust incorrect outputs.","core_discovery":"LinkQ is posed as a testbed: an LLM-assisted knowledge-graph exploration system whose five visual mechanisms are intended to make the query-generation pipeline inspectable. The central discovery is that this transparency may backfire. The paper reports that users, even KG experts, tended to overtrust LinkQ's outputs when the LLM was wrong, citing the state diagram and the query structure graph as reasons for confidence. Incorrect answers were sometimes rationalized as plausible rather than rejected. The authors attribute this overconfidence to the added visual transparency of the natural-language interface, while also observing that users adopted distinct workflows depending on their knowledge-graph and LLM experience, so a single interface design will not serve all users.","pith_inferences":["Editorial inference — The study had no no-visualization baseline, so the overtrust could instead come from the LLM's textual authority; a chat-only control condition would separate these causes.","Editorial inference — A quantitative follow-up that toggles each visualization off and measures acceptance of seeded wrong answers could rank the components by their trust effect; participant quotes suggest the LLM-KG state diagram is the most likely driver.","Editorial inference — If visual plausibility drives false confidence, then visualizations that actively challenge the answer—for example, by showing alternative queries or explaining empty results—may reduce overtrust more than neutral transparency panels; the authors gesture at this direction but do not test it."],"forward_implications":["If the finding holds, transparency visualizations in LLM-assisted tools are trust-shaping features, not neutral explainability aids, and should be designed with the possibility of false confidence in mind.","Users with different knowledge-graph and LLM experience will need different supports: query-graph inspection for those who can read queries, chat and entity-relation context for those who cannot, and source-linking for skeptical users.","The results graph view went unused while tabular results dominated, so graph-style output needs a concrete exploratory purpose rather than being included for visual appeal.","Because participants tackled open-ended tasks by asking narrow, deliberate sub-questions, LLM-assisted exploration may reduce overall data coverage compared to unassisted exploration."],"supporting_citations":[{"why":"It documents that users over-rely on AI even when errors are present, the background phenomenon that LinkQ's overtrust result echoes.","marker":"[11]"},{"why":"It supplies the trust-in-visualization framework the authors use to interpret how visual transparency shifts user confidence.","marker":"[13]"},{"why":"It provides evidence that AI-guided visual analytics can bias trust and data exploration, a direct comparator for the LinkQ findings.","marker":"[24]"},{"why":"It shows recent evidence that conversational AI explanations can increase overreliance, which the paper aligns with its own overtrust observations.","marker":"[26]"},{"why":"It establishes that plausible-sounding AI justifications make users accept incorrect information, the mechanism invoked for users' rationalizing wrong answers.","marker":"[31]"},{"why":"It is the earlier LinkQ system paper that defines the tool used as the study's testbed.","marker":"[41]"},{"why":"It shows charts can mislead through framing rather than visual error, supporting the idea that well-intentioned visuals can produce false belief.","marker":"[45]"},{"why":"It surveys LLM hallucination as an inherent limitation, motivating why human oversight and trust calibration are needed.","marker":"[58]"}],"fun_headline_variants":["Helpful visuals make even experts trust bad LLM answers","KG experts overtrust wrong LLM results when visuals look clear","Overtrust in LLM outputs grows with visualization clarity","Visual clarity backfires: users trust false LLM answers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the visualizations caused overtrust rests on the assumption that this effect can be separated from the general authority of the LLM and from users' pre-existing trust dispositions; the study had no condition without or with altered visualizations to test that separation.","fun_headline_variants_meta":{"raw":{"variants":["Helpful visuals make even experts trust bad LLM answers","KG experts overtrust wrong LLM results when visuals look clear","Overtrust in LLM outputs grows with visualization clarity","Visual clarity backfires: users trust false LLM answers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000613,"raw_usage":{"total_tokens":2865,"prompt_tokens":977,"completion_tokens":1888,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":1820}},"tokens_in":593,"tokens_out":1888,"duration_ms":12932,"temperature":1.0,"reasoning_tokens":1820,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:27:38.657592+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Present matched groups of participants with the same deliberately wrong LinkQ answers, one group using the full visual interface and one using a chat-only version with identical LLM text but no state diagram, query graph, or entity-relation table; if wrong-answer acceptance rates are equal, the visualizations are not the cause of overtrust.","supporting_citations":[{"cited_title":"Buçinca, M","cited_arxiv_id":null,"evidence_quote":"It documents that users over-rely on AI even when errors are present, the background phenomenon that LinkQ's overtrust result echoes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the trust-in-visualization framework the authors use to interpret how visual transparency shifts user confidence."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides evidence that AI-guided visual analytics can bias trust and data exploration, a direct comparator for the LinkQ findings."},{"cited_title":"Is Conversational XAI All You Need? Human-AI Decision Making With a Conversational XAI Assistant","cited_arxiv_id":"2501.17546","evidence_quote":"It shows recent evidence that conversational AI explanations can increase overreliance, which the paper aligns with its own overtrust observations."},{"cited_title":"Jacovi, J","cited_arxiv_id":null,"evidence_quote":"It establishes that plausible-sounding AI justifications make users accept incorrect information, the mechanism invoked for users' rationalizing wrong answers."}],"review_version":1}