{"id":"14bad23b-ec67-4ca3-91eb-d5819b7ae9e6","arxiv_id":"2608.09698","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"VeriForge demonstrates that proactive, source-traceable highlighting of potential knowledge gaps can help fiction writers discover and integrate unfamiliar domain facts during drafting.","lead":"VeriForge is an AI writing tool that detects domain knowledge gaps fiction writers cannot see themselves, flagging unfamiliar concepts and pairing them with source-anchored fact cards as the author drafts. A 12-writer study found it helped authors notice overlooked gaps and produced passages that blind expert raters judged as showing stronger domain competence.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No objective fact-checking of highlights or final passages; 'domain grounding' is measured only by perceived competence, so the knowledge-gap claim is not yet established.","rationale":"The reader identified the precision of proactive highlighting as the weakest assumption. I agree it is central, but the more load-bearing and decision-relevant unvalidated link sits one step downstream: even if highlight selection is perfect, the LLM-generated card descriptions could be wrong, and the study never checks whether authors' final passages are factually grounded. The paper is honest about this—Section 7.3 calls the RQ4 ratings 'initial evidence,' and the abstract carefully says 'perceived by expert raters'—but the title and contribution 3 claim 'improves domain grounding,' which is stronger than what was measured. The deception experiment strengthens the concern: 11/12 participants did not inspect the source passage, so a wrong AI description would be integrated uncritically. This is not an internal inconsistency; the reported statistics and qualitative data are consistent with what was measured. It is a gap between the measured construct (perception) and the asserted construct (actual mitigation of latent knowledge gaps). Positive evidence that should be credited: the strengthened baseline sharing retrieval infrastructure is a good design choice, FDR correction is applied transparently, delayed recall is an objective behavioral measure, and the deception result is candidly reported. Because the paper is explicitly preliminary and the reader already conditioned acceptance, this concern does not change the verdict; it sharpens the condition: an objective factual-grounding evaluation (or at least a highlight-precision/description-accuracy audit) should be required before the full framing is accepted.","tokens_in":21140,"tokens_out":7392,"duration_ms":72378,"concrete_test":"Have two domain experts (e.g., a HEMA practitioner and a TCMA practitioner) independently fact-check all 24 anonymized story snippets against the corresponding uploaded source corpora, coding each for factual errors, unsupported domain claims, and grounding accuracy, and compare error rates between the VeriForge and baseline conditions. If VeriForge passages do not contain significantly fewer factual errors (or materially higher factual grounding) than baseline, the 'stronger domain grounding' component of the central claim fails, even if perceived competence and self-reported gap awareness remain higher.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that VeriForge surfaces and helps authors integrate genuine domain facts. The study never measures whether that is true. Section 4.2.1 defines proactive highlighting as a 'low-cost hypothesis about knowledge value,' but no precision, recall, or false-positive analysis is reported for the Alignment Engine. Knowledge Card descriptions are LLM-generated, yet their factual accuracy against the source passage is not evaluated anywhere. RQ4 (Section 5.5.3) operationalizes 'domain grounding' solely through two published authors' holistic ratings of 'perceived authorial domain competence'; the raters are not identified as HEMA/TCMA domain experts and were not asked to verify facts against the uploaded source corpora. The deception experiment (Section 6.3) shows that 11/12 participants accepted a mismatched source card without inspection, so users are unlikely to catch inaccurate AI descriptions. Consequently, the observed increases in queries, canvas nodes, delayed recall, and perceived competence are consistent with a system that makes passages sound more grounded without actually filling the underlying knowledge gaps. The load-bearing assumption—that highlighted knowledge is correct and genuinely gap-filling—is untested.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VeriForge, a mixed-initiative writing environment for fiction authors working in unfamiliar domains. The system combines proactive inline highlighting that flags candidate knowledge gaps, dual-stream querying that pairs conversational responses with source-anchored Knowledge Cards, and a spatial Knowledge Canvas for organizing discovered facts, all backed by a graph-based retrieval-augmented generation pipeline over user-uploaded source material. The design is motivated by formative interviews with nine fiction writers and evaluated in a within-subjects user study (N=12) against a strengthened baseline that shares the same retrieval backend but omits the proactive and canvas-integration mechanisms. The authors report that VeriForge increases blind-spot awareness, query volume, canvas use, creativity support, delayed recall, and one of three expert-rated prose dimensions (perceived authorial domain competence), while preserving perceived authorial autonomy.","tokens_in":21405,"tokens_out":4103,"duration_ms":40735,"significance":"If the findings hold, VeriForge addresses a genuinely underexplored problem: helping writers discover domain knowledge they cannot self-diagnose, rather than only retrieving answers to explicit queries. The manuscript has several real strengths: the formative study grounds the design in observed writer behavior; the within-subjects design with a strengthened baseline is a thoughtful attempt to isolate interaction design from retrieval quality; FDR correction is applied within predefined families; the deception experiment directly probes epistemic trust; and the delayed-recall analysis is a useful exploratory addition. The reported effect sizes are large and consistent across several behavioral and self-report measures. However, the central construct—mitigating latent knowledge gaps—is never validated against ground truth: highlight precision, Knowledge Card factual accuracy, and the factual correctness of final passages are not measured. Given the deception result showing that users do not inspect provided sources, the observed benefits could in principle stem from plausible but inaccurate scaffolding.","major_comments":[{"comment":"The central mechanism of VeriForge is proactive highlighting, defined in Section 4.2.1 as a 'low-cost hypothesis about knowledge value.' The paper reports no precision, recall, or false-positive analysis for the Alignment Engine's highlight suggestions, and no factual-accuracy evaluation of the LLM-generated Knowledge Card descriptions against their quoted source passages. The deception experiment in Section 6.3 shows that 11 of 12 participants accepted a card whose source passage was replaced with unrelated text, so users are unlikely to catch inaccurate AI-generated descriptions. Without an audit showing that highlights and cards correspond to genuine, source-supported knowledge, the claim that VeriForge 'mitigates latent knowledge gaps' is not yet established; the evidence supports only the weaker claim that it changes perceived gap awareness and exploration behavior. Please add an objective evaluation (e.g., expert or LLM-assisted fact-checking of a sample of highlights and cards against the source corpora, with error rates) or explicitly restrict the contribution to perceived gap recognition.","section":"Section 4.2.1 and Section 6.3"},{"comment":"RQ4 is the only outcome measure tied to the 'domain grounding' claim, but it is operationalized solely through two raters' holistic judgments of 'perceived authorial domain competence.' The raters are described as published authors, not as HEMA or TCMA domain experts, and they were not asked to verify any factual claims in the passages against the uploaded source corpora. The manuscript therefore never measures whether the knowledge gaps were actually filled correctly. Given that the deception study shows readers do not verify sources, the observed difference in perceived competence could reflect surface-level use of domain terminology rather than accurate grounding. I recommend adding a fact-based outcome measure—for example, the proportion of domain-specific claims in the final passages that are accurate according to the provided sources, or a count of factual errors—or, failing that, changing the abstract's 'stronger domain grounding' to 'higher perceived authorial domain competence.'","section":"Section 5.5.3 and Section 6.4"},{"comment":"The expert-rating evidence is thinner than the abstract implies. Of the three rated dimensions, only 'perceived authorial domain competence' survives FDR correction (q = .015); natural integration (q = .065) and specificity (q = .065) are positive trends that do not meet the corrected threshold, and the composite score reported in Section 6.4 is explicitly uncorrected. The single surviving dimension is also the most perceptual and least tied to factual accuracy. This is not a fatal flaw for a preliminary study, but it should be stated clearly in the abstract and conclusion so that readers do not infer broad support for 'stronger domain grounding' across all quality dimensions.","section":"Section 5.5.4, Section 6.4, and Figure 5"}],"minor_comments":[{"comment":"The term 'expert raters' in the abstract is stronger than the description in Section 5.5.3, where the raters are 'two published authors who were not study participants' and not necessarily domain experts. Please align the terminology throughout.","section":"Section 5.5.3"},{"comment":"The caption notes that the expert composite is uncorrected, but the main text should also prominently state that natural integration and specificity did not survive FDR correction, since these are reported in the same paragraph as the composite result.","section":"Figure 5 caption and Section 6.4"},{"comment":"The delayed-recall analyses are described as exploratory and are not FDR-corrected, yet they are presented with strong effect sizes and p-values. Please clearly label these results as exploratory in the figures or tables, not only in the prose.","section":"Section 5.5.2 and Section 6.2"},{"comment":"For the perceived authorial autonomy result (M = 6.17 vs. M = 6.17, p = 1.0), please report the number of nonzero differences and the distribution of signed ranks; with N = 12, a p-value of exactly 1.0 can arise from mostly tied responses, and this information is needed to interpret the null result.","section":"Section 6.3"},{"comment":"The implementation section names Qwen3.5-Flash and BAAI/bge-m3 but does not report the entity-extraction or alignment prompt templates or any examples of highlight failures. Including representative positive and negative highlight examples in the supplemental material would help readers calibrate the reliability of the proactive mechanism.","section":"Section 4.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-motivated systems/HCI contribution with an unusually honest limitations section, and the strengthened baseline is a methodological improvement over many similar studies. The main obstacle to acceptance is that the paper's headline claim—mitigating latent knowledge gaps—is not directly measured; the evaluation establishes perceived gap awareness and perceived grounding, but the deception result makes the absence of objective fact-checking especially consequential. The authors could likely address this with a focused audit of highlight/card accuracy and a fact-based passage-groundedness measure, or by carefully restricting the claims in the title and abstract. I would not recommend rejection, but the revision must either supply the missing objective validation or reframe the contribution accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read VeriForge. It's a solid systems paper, better than most UIST submissions, and worth engaging seriously. The genuinely new piece is the interaction pattern: proactive inline highlighting of candidate knowledge gaps paired with a dual-stream source-anchored card layout and a spatial canvas, all justified by a discovery/synthesis split under IOED. That split is well argued and connects to a real cognitive failure. I also credit the study design—the strengthened baseline that shares the same GraphRAG backend is exactly the right move to isolate interaction from retrieval quality, and the authors were transparent enough to report which outcomes survived FDR correction.\n\nWhat I'd want a referee to probe hardest is the accuracy of the proactive content. The Alignment Engine and LLM-generated Knowledge Cards are never evaluated for precision or factual correctness. Section 4.2.1 explicitly calls a highlight a 'low-cost hypothesis,' but the load-bearing claim is that these hypotheses point at real gaps. The stress-test note has it right: RQ4's 'domain grounding' is measured only by two raters' perceived authorial competence, not by checking facts against source corpora. Combined with the deception experiment—11/12 participants accepted a mismatched source card—this leaves the central claim of actually filling knowledge gaps partly supported at best. The observed effects are consistent with the system increasing perceived authority and exploration without verifiably correcting domain errors.\n\nOther soft spots are minor or honestly disclosed: N=12, moderate ICCs, no code/data release, and only one of three expert-rater dimensions surviving correction. I don't think the authors are hiding anything; the provenance paradox discussion is unusually candid. But the knowledge-gap mechanism itself needs a technical evaluation before the full framing is established.\n\nWho is this for? HCI researchers working on AI-assisted writing, mixed-initiative systems, and creativity support tools. It deserves serious peer review; I'd send it out. For publication I'd ask for a highlight-hit-rate analysis against expert annotations, an accuracy check on Knowledge Card descriptions, and ideally a small fact-count verification of the final passages.","headline":"A thoughtful mixed-initiative writing study with an honest baseline and one honest design flaw: the knowledge the system surfaces is never checked for factual accuracy, so the headline gap-filling claim rests on perception, not verification.","tokens_in":21865,"tokens_out":1596,"would_cite":true,"duration_ms":15967,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Proactive, source-traceable scaffolding helps fiction writers discover and fill domain-knowledge gaps they cannot name on their own.","keywords":["AI-assisted writing","mixed-initiative interaction","narrative verisimilitude","knowledge gaps","retrieval-augmented generation","creativity support tools","provenance","user study"],"falsifier":"Run VeriForge on drafts that have been seeded with known domain errors, such as a still-bleeding corpse described as decomposing or an anachronistic vaccine rollout. If the proactive highlights fail to flag a large share of seeded errors, or flag a large share of correct statements, the claim that the system reveals latent gaps instead of injecting noise would be refuted.","tokens_in":1505,"feed_emoji":"✍️","tokens_out":1735,"duration_ms":57762,"temperature":0.7,"pith_summary":"The paper argues that the most damaging gaps in a fiction writer's domain knowledge are those authors cannot diagnose themselves, an instance of the illusion of explanatory depth. It presents VeriForge, a mixed-initiative writing environment that splits cognitive labor: the system proactively flags possible knowledge gaps and supplies source-traceable facts, while the author keeps full control over turning those facts into prose. In a within-subjects study with 12 writers, the authors report that VeriForge helped participants notice overlooked gaps, prompted more queries and canvas use, raised perceived creativity support, and produced passages that blind expert raters judged stronger in perceived domain competence. The paper positions these results as preliminary evidence, not a settled proof.","feed_headline":"AI scaffold reveals fiction writers' hidden knowledge gaps","feed_subtitle":"Inline highlights plus source-traceable cards helped 12 writers spot blind spots and draft scenes experts rated as more domain-competent.","key_machinery":"The load-bearing mechanism is the division of cognitive labor. VeriForge's front end couples three interactions: proactive inline highlights triggered after a five-second pause when an Alignment Engine finds a source-backed term whose graph neighborhood suggests useful knowledge; dual-stream queries that present the same retrieval as both a conversational text stream and source-anchored Knowledge Cards; and a spatial Knowledge Canvas whose nodes carry provenance and whose topology feeds back into the Semantic Frame for later retrievals. Behind them, a graph-based retrieval-augmented generation pipeline stores domain facts as nodes and edges with exact source passages bound at ingestion, so every surfaced fact remains traceable to its primary source. The Alignment Engine is the piece that decides what counts as a candidate gap, using a hand-coded relational-complexity taxonomy rather than error detection.","core_discovery":"The central claim is that proactive, flow-aware, source-traceable scaffolding can make authors aware of latent domain-knowledge gaps that reactive tools cannot surface, without taking narrative authorship away. The authors ground this in the knowledge-transforming model of composition: writing well in an unfamiliar domain requires shuttling between factual content and narrative rhetoric, and the illusion of explanatory depth blocks that shuttle precisely where it matters most. VeriForge operationalizes a discovery-synthesis split, in which the system takes initiative over finding and structuring external facts while the author takes initiative over composing, and the study's pattern of results, higher blind-spot alerting, more exploration, higher perceived creativity support, higher expert-rated domain competence, and unchanged autonomy, is offered as preliminary evidence that the split works.","pith_inferences":["If the discovery-synthesis split is the active ingredient, the same scaffolding logic should transfer to journalism, documentary filmmaking, legal drafting, and clinical writing, wherever the output's value depends on a human voice interpreting expert material; the paper gestures at this transfer but does not test it.","The unmeasured precision of highlight selection is the critical knob: a risk-adaptive design that raises epistemic friction when consequences are higher could convert the provenance-as-authority failure into a feature while preserving flow in low-stakes writing.","A direct testable extension would replace the AI description rather than the source passage in the deception experiment, isolating whether authors can detect fabricated domain claims; the paper notes this is not tested.","The retention-rate gains suggest that writing with proactive factual scaffolding may function like writing-to-learn, but the 30-minute cold-start task cannot show whether the effect compounds over a full manuscript; longitudinal drafts would settle that."],"forward_implications":["Authors using VeriForge in the study rated blind-spot alerting far higher than they did in a reactive baseline sharing the same retrieval backend, and they issued more queries, so proactive cues changed search behavior rather than merely improving satisfaction.","Blind expert raters scored VeriForge passages higher on perceived authorial domain competence, with positive but uncorrected trends on natural integration and specificity, suggesting the paradigm can affect cold-start prose quality.","Participants' perceived authorial autonomy was identical across conditions, supporting the claim that a discovery-synthesis split preserves agency even when the system takes initiative over knowledge discovery.","One-week delayed recall showed that participants retained more domain terms and a higher proportion of the terms they had used under VeriForge, consistent with deeper encoding rather than mere exposure.","Only 1 of 12 participants noticed a deliberately mismatched source passage, so visible provenance can act as an authority cue that suppresses verification even when the underlying facts are trustworthy."],"supporting_citations":[{"why":"Supplies the illusion of explanatory depth, the cognitive mechanism that makes authors blind to the knowledge gaps the system targets.","marker":"[55]"},{"why":"Supplies the knowledge-transforming model of composition that motivates the discovery-synthesis split.","marker":"[11]"},{"why":"Provides the mixed-initiative principles used to justify giving the system discovery initiative and the author synthesis initiative.","marker":"[28]"},{"why":"Supplies the Graph RAG approach the back-end adapts for relational, provenance-bound retrieval.","marker":"[21]"},{"why":"Represents the reactive conversational paradigm the system is designed to overcome and the baseline workflow mimics.","marker":"[47]"},{"why":"Defines verisimilitude and worldbuilding standards that set the goal for domain grounding.","marker":"[67]"},{"why":"Provides the sensemaking scale used to measure how well the canvas supports relational knowledge building.","marker":"[4]"},{"why":"Underlies the dual-process interpretation of why users accepted a mismatched source card without verification.","marker":"[50]"},{"why":"Explains how source labels and structured formatting function as credibility heuristics, the provenance-as-authority effect observed in the deception test.","marker":"[59]"}],"fun_headline_variants":["VeriForge exposes the fiction gaps you don't know you have","Proactive flags uncover latent gaps in fiction drafting","AI reveals blind spots, authors stay in charge","Tool+cards help 12 writers draft more domain-accurate scenes"],"cache_read_input_tokens":24064,"weakest_assumption_plain":"The load-bearing premise is that the proactive highlights point at real, useful knowledge gaps most of the time; the paper does not measure how often they are right, so if highlights are frequently irrelevant or misleading the reported benefits would not survive.","fun_headline_variants_meta":{"raw":{"variants":["VeriForge exposes the fiction gaps you don't know you have","Proactive flags uncover latent gaps in fiction drafting","AI reveals blind spots, authors stay in charge","Tool+cards help 12 writers draft more domain-accurate scenes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001436,"raw_usage":{"total_tokens":5805,"prompt_tokens":975,"completion_tokens":4830,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":4761}},"tokens_in":591,"tokens_out":4830,"duration_ms":32002,"temperature":1.0,"reasoning_tokens":4761,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:29:33.666775+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run VeriForge on drafts that have been seeded with known domain errors, such as a still-bleeding corpse described as decomposing or an anachronistic vaccine rollout. If the proactive highlights fail to flag a large share of seeded errors, or flag a large share of correct statements, the claim that the system reveals latent gaps instead of injecting noise would be refuted.","supporting_citations":[{"cited_title":"1987.The Psychology of Written Compo- sition","cited_arxiv_id":null,"evidence_quote":"Supplies the knowledge-transforming model of composition that motivates the discovery-synthesis split."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the reactive conversational paradigm the system is designed to overcome and the baseline workflow mimics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the sensemaking scale used to measure how well the canvas supports relational knowledge building."},{"cited_title":"Petty and John T","cited_arxiv_id":null,"evidence_quote":"Underlies the dual-process interpretation of why users accepted a mismatched source card without verification."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Explains how source labels and structured formatting function as credibility heuristics, the provenance-as-authority effect observed in the deception test."}],"review_version":1}