{"id":"6d2ba29c-3cb5-4934-bfde-160e331bce56","arxiv_id":"2506.03052","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Layering structured visual representations over unstructured chat feedback helps novice designers notice, revisit, and learn the design principles behind the feedback.","lead":"Feedstack is a new chatbot interface that layers bookmarks, chapters, and highlights on top of ordinary feedback conversations, helping users find and understand the design principles behind the feedback. This design paper is for researchers and builders of conversational AI who want to move chat beyond simple back-and-forth dialogue.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The user study tested a pre-populated conversation, so the real-time LLM chapter-generation pipeline—the core of the 'shared representation' claim—was never exercised; if that pipeline is inaccurate, the scaffolding could mislead novices.","rationale":"The reader's weakest_assumption already identified the untested LLM chapter pipeline, and my reading confirms this is the most load-bearing concern. The design-probe framing in Section 6.3 explicitly disclaims conclusive evaluation, so the absence of live-pipeline testing does not warrant outright rejection; it does, however, mean the central 'shared representation' claim is supported only by a Wizard-of-Oz-style simulation. The proposed concrete test would directly measure whether the automatic extraction is reliable enough to produce the claimed scaffolding. If it passes, the existing CONDITIONAL verdict could be upgraded; if it fails, the value claim for the live system would need substantial revision. I therefore see no reason to change the reader's CONDITIONAL verdict.","tokens_in":9128,"tokens_out":2996,"duration_ms":32584,"concrete_test":"Run the live Feedstack pipeline on the eight formative-study conversations and on, say, 10 new design-feedback dialogues. Have two design experts independently label every mention of each of the five design principles; then compare the LLM-generated chapters/bookmarks against these labels using precision, recall, and F1, and rate the correctness of the generated principle definitions and 'Relation to Your Design' text. If recall or correctness is low (e.g., F1 < 0.8 or >25% of definitions rated inaccurate), the study's positive results are not evidence for the real-time chapter-generation claim and an updated user study with live generation is required.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 states 'Chapters are generated in real time using a large language model to analyze the conversation,' and the opacity feature is said to signal 'to users and the chatbot which topics are of interest.' But Section 5.1 says the user study 'populated the conversation as text in the prototype, enabling users to progress through the design feedback conversation'—i.e., participants read a pre-authored script rather than interacting with a live, LLM-generated set of chapters. All reported benefits (e.g., P6 noticing Consistency, P7 noticing Balance instances) are thus about hand-constructed or precomputed content. The central claim that layered structures 'surface user intent and reveal underlying design principles' depends on the automatic extraction being accurate; if the LLM misassigns principles or writes plausible but wrong definitions in the Learning Materials, the scaffolding may reinforce misconceptions rather than correct them. There is also no evidence in the paper that the chatbot itself consumed the chapter state during the conversation, so 'shared representation' is only partially demonstrated. The paper is honest about its design-probe status (Section 6.3), but the unvalidated live pipeline is the critical load-bearing element.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Feedstack, a speculative conversational user interface that layers structured representations—bookmarks, chapters, highlights, principle toggles—on top of an unstructured chatbot feedback conversation. The authors report two formative user studies (n=8 each) with novice designers, the first observing interactions with ChatGPT and the second a think-aloud study with a pre-populated Feedstack conversation. The central claim is that these layered structures act as a shared representation between user and AI, surfacing implicit design principles, supporting exploration and reflection, and helping novices connect feedback instances to broader design principles. The paper is explicitly positioned as a research-through-design artifact and design probe rather than a controlled evaluation.","tokens_in":9376,"tokens_out":3402,"duration_ms":41714,"significance":"If the central claim holds, Feedstack contributes a concrete interaction pattern for educational CUIs: adding non-linear structure around a linear chat without destroying conversational flow. The paper's strengths are its grounding in established design literature (variation theory, dual-coding theory, shared representations, information scent), its clear and honest framing as a design probe, its two complementary small studies, and its use of participant quotes to illustrate design insights. The prototype is functional and the design rationale is reproducible. However, the significance is limited by the lack of a live-system evaluation and by the absence of any comparison condition, so the contribution is best understood as a design exploration that generates hypotheses rather than as evidence that the real-time pipeline benefits learning.","major_comments":[{"comment":"The central claim that the layered structures constitute a shared representation that 'surfaces user intent and reveals underlying design principles' depends on the real-time LLM chapter-generation pipeline described in Section 4.2, but the user study in Section 5.1 'populated the conversation as text in the prototype' rather than exercising live generation. Consequently, none of the participant quotes in Section 5.2, such as P6 noticing Consistency or P7 noticing Balance instances, provide evidence about the accuracy, timing, or helpfulness of automatically generated chapters, learning materials, or opacity updates. If the LLM misassigns principles or generates plausible but incorrect definitions, the scaffolding could mislead novices rather than support them. The authors should either test the live pipeline with real conversations or explicitly limit the paper's contribution to the interaction design of pre-structured content, and temper the abstract and conclusion claims accordingly.","section":"4.2, 5.1, 6.3"},{"comment":"The study is an eight-participant think-aloud session without a comparison condition, and all reported benefits are self-reported perceptions rather than observed learning or changes in design behavior. For example, P2's 'teach them to fish' quote and P8's statement about future use are intentions, not evidence of improved understanding. The paper already frames itself as research-through-design, which is acceptable for generating hypotheses, but the conclusion that the tool 'helps novice designers become more aware' (Section 6.1.1) goes beyond what this design can support. I recommend replacing knowledge-claim language with design-insight language, e.g., 'participants reported that the affordances directed their attention to previously unnoticed principles.'","section":"5.1, 5.2, 6.1.1"},{"comment":"The five-principle taxonomy (Accessibility, Consistency, Contrast, Balance, Alignment and Spacing) is presented as a fixed, sufficient set for the target feedback conversations, but no analysis is provided to show that this taxonomy covers the actual feedback utterances in the two studies. Since chapter opacity, bookmarks, and highlights are defined relative to these principles, an incomplete or mismatched taxonomy would systematically hide relevant feedback topics. The paper should report how the taxonomy was derived and whether any utterances in the formative study fell outside it, or acknowledge this as a limitation and an avenue for future work.","section":"4.2, 5.1"}],"minor_comments":[{"comment":"In the introduction, the sentence 'These ‘shared representations also serve to scaffold...' has an unclosed quotation mark and should read 'These shared representations also serve...'.","section":"1"},{"comment":"The phrase 'Many students now turning to general purpose chatbots' is missing an auxiliary verb and should read 'Many students are now turning...'.","section":"2.1"},{"comment":"The phrase 'conversations;for example' is missing a space after the semicolon; it should read 'conversations; for example'.","section":"2.2"},{"comment":"The caption labels panels A-F, and the body text refers to 'the design panel A'; considering the chat panel is also a design element, renaming A to 'artifact panel' would reduce potential confusion.","section":"Figure 1 caption"},{"comment":"Figure 3 is introduced after the discussion of Emerging Topics, Conversational Cues, and Referencing Conversation, but the text does not explicitly link the labeled elements L and M in the figure to these feature names; a brief cross-reference in the text would improve clarity.","section":"6.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-written design probe with honest limitations, and the authors clearly state that they are not presenting a conclusive evaluation. My recommendation is driven by the gap between the abstract's central claim about a 'shared representation' that surfaces user intent and the fact that the LLM-driven chapter generation was never exercised in the reported user study. The paper could be acceptable for a workshop venue as is, but for archival publication the authors should either validate the live pipeline or explicitly narrow the contribution to the interaction design of pre-structured content."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Feedstack is a small, honest research-through-design paper, and you should read it as exactly that. The contribution is a concrete design artifact: layering bookmarks, chapters, and highlights onto a live design-feedback chat, with the rationale that these structures externalize implicit design principles. That specific integration is new relative to the earlier systems they cite (Sensecape, CausalMapper, Graphologue, Decipher, Voyant), even though each individual affordance has clear precedents. This is not a new capability or a discovered law; it is a reusable design direction for conversational feedback tools. The paper is appropriately pitched as such.\n\nI want to give credit where it is earned. The design rationale is grounded in existing literature (variation theory, dual-coding, shared representations) and the formative study that generated the design goals is a sensible first step. The second study is small but the participant quotes do support the stated claims about awareness and reflection. The authors are also upfront that this is a design probe, not a conclusive evaluation, and they name the lack of comparison with commercial chatbots in Section 6.3. The citation pattern looks fine; the related work is on-target.\n\nNow the soft spots, in proportion. The biggest one is the one the stress-test flags: Section 4.2 says chapters are generated in real time by an LLM, but the study in Section 5.1 populated the conversation as pre-authored text. So the central mechanism behind the 'shared representation' claim was never actually exercised. The benefits you see in the results come from hand-constructed or precomputed content, not from a live pipeline that detects principles and writes definitions on the fly. If that pipeline mislabels or generates plausible-but-wrong material, the scaffolding could actively mislead novices. The authors do not test this, and they do not claim to. A second soft spot: there is no baseline condition, so you cannot tell whether the benefits come from the layering or from simply showing design principles in a side panel. The n=8 think-aloud data is fine for formative insight, but it does not support strong claims.\n\nNone of this is fatal, because the paper frames itself as an early exploration. The useful takeaway is that the design idea is worth engaging with, but the live LLM pipeline is the load-bearing piece that would need to be evaluated before this becomes more than a promising prototype.\n\nWho is this for? Researchers working on conversational UI, design feedback, and learning tools. I would bring it to a reading group as an example of disciplined design-probe work. It deserves a serious referee: it is a coherent, well-scoped CUI short paper that would benefit from constructive review rather than a desk reject. I would not cite it as evidence of effectiveness, but I would cite it as a design pattern.","headline":"A modest but honest design-probe paper: the layered chat affordances are the contribution, and the evaluation is formative—treat it as a design exploration, not a validated system claim.","tokens_in":9891,"tokens_out":1055,"would_cite":true,"duration_ms":14556,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Layering structure onto a design-feedback chat surfaces implicit design principles for novice designers.","keywords":["conversational user interface","design feedback","design principles","shared representation","scaffolding","reflection","novice designers","research through design"],"falsifier":"Have design experts annotate a set of real feedback conversations for the design principle each turn discusses, then run Feedstack's real-time chapter generator over the same conversations; if the generated chapters systematically miss or mislabel principles, the shared representation would mislead users rather than scaffold them.","tokens_in":8937,"feed_emoji":"🧩","tokens_out":5803,"duration_ms":66346,"temperature":0.7,"pith_summary":"The paper introduces Feedstack, a conversational interface that layers bookmarks, chapters, and highlights onto an ordinary design-feedback chat with an AI. Its central claim is that these layered structures form a shared representation of the conversation that makes implicit design principles explicit, helping novice designers notice and reflect on what the feedback is really about. Two formative studies, each with eight beginners, produced reports of participants discovering principles they had overlooked and revisiting the conversation after it ended to apply the feedback. The paper frames this as an early design probe, not a conclusive evaluation.","feed_headline":"Structured chat layers reveal hidden design principles","feed_subtitle":"Bookmarks, chapters, and highlights turn a design-feedback chat into a navigable surface for novice learners.","key_machinery":"The load-bearing mechanism is the set of layered affordances wrapped around the chat transcript. Bookmarks place visible markers on the conversation's scrub bar wherever a design principle is discussed, so users can jump between instances. Chapters are LLM-generated accordions, each tied to a design principle, containing a definition, an explanation of how the principle applies to the user's design, and key terms; their opacity increases as the principle is discussed more. Highlights color key terms in the chat, and principle toggles switch them on and off. Together these layers are the system's shared representation of the conversation, giving the user and the AI common ground.","core_discovery":"Feedstack's central claim is that a feedback conversation does not have to live only as a linear transcript. By bookmarking where each design principle is discussed, generating chapter summaries of those principles as the conversation unfolds, and highlighting key terms in place, the interface externalizes the tacit structure of a critique. Users can jump from a bookmark to an earlier 'Balance' moment, watch the Balance chapter expand beside the chat, and see how frequently each principle has come up through the chapter's opacity. That externalization is what lets a novice zoom out from turn-by-turn dialogue and grasp the larger themes, including principles that have not yet been discussed.","pith_inferences":["The step most worth stress-testing is the real-time chapter generator: if the language model mislabels which principle a turn discusses, the whole shared representation would teach the wrong structure to learners.","The same layered-representation idea could transfer to other apprenticeship conversations, such as code review or scientific peer review, where the underlying principles are tacit for novices.","Opacity could be reused as a live coverage dashboard, making it easy to test whether a conversation has become unbalanced toward a few principles.","A natural follow-up experiment would measure learning, not just preference: give two groups of novices the same feedback, with and without the layered panel, and compare how accurately they can recall and apply the principles."],"forward_implications":["Novice designers can become aware of principles that were never named outright in the feedback, such as consistency or accessibility.","After the chat ends, the transcript becomes a navigable reference, so users can return to specific feedback instances while revising their designs.","Seeing a principle discussed in multiple bookmarked places supports comparing how the same principle plays out in different parts of a design.","Opacity and suggested emerging topics can point users toward design principles the conversation has so far neglected.","These benefits can appear without resorting to a rigid, wizard-driven dialogue, preserving the open-ended feel of a real critique."],"supporting_citations":[{"why":"frames the prototype as a design probe, shaping how the two studies should be read.","marker":"[3]"},{"why":"supplies the dual-coding rationale behind highlighting key terms in the transcript.","marker":"[7]"},{"why":"motivates the need to connect feedback instances to high-level principles for learning.","marker":"[12]"},{"why":"grounds the shared-representation rationale for Chapters in mixed-initiative systems.","marker":"[13]"},{"why":"provides the five core design principles the chapter structure is built around.","marker":"[23]"},{"why":"supports the reflection goal that bookmarks serve by enabling revisiting and review.","marker":"[29]"},{"why":"defines the exploratory design-probe approach the two studies follow.","marker":"[40]"}],"fun_headline_variants":["Layered chat reveals design principles novices miss","Bookmarks and chapters turn chat into a design map","Externalize critique structure with Feedstack's layers","Layered feedback chat for smarter design critiques"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The scaffolding effect rests on the AI reliably recognizing when a design principle is being discussed and generating accurate chapter content, a step the paper never tested with a live, automatically generated conversation.","fun_headline_variants_meta":{"raw":{"variants":["Layered chat reveals design principles novices miss","Bookmarks and chapters turn chat into a design map","Externalize critique structure with Feedstack's layers","Layered feedback chat for smarter design critiques"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000491,"raw_usage":{"total_tokens":2355,"prompt_tokens":825,"completion_tokens":1530,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":1471}},"tokens_in":441,"tokens_out":1530,"duration_ms":11282,"temperature":1.0,"reasoning_tokens":1471,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:10:02.557768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have design experts annotate a set of real feedback conversations for the design principle each turn discusses, then run Feedstack's real-time chapter generator over the same conversations; if the generated chapters systematically miss or mislabel principles, the shared representation would mislead users rather than scaffold them.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"frames the prototype as a design probe, shaping how the two studies should be read."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the dual-coding rationale behind highlighting key terms in the transcript."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"motivates the need to connect feedback instances to high-level principles for learning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"grounds the shared-representation rationale for Chapters in mixed-initiative systems."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the five core design principles the chapter structure is built around."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supports the reflection goal that bookmarks serve by enabling revisiting and review."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the exploratory design-probe approach the two studies follow."}],"review_version":1}