Pith. sign in

REVIEW 3 major objections 5 minor 39 references

Feedstack: Layering Structured Representations over Unstructured Feedback to Scaffold Human AI Conversation

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Layering structure onto a design-feedback chat surfaces implicit design principles for novice designers.

desk verdict A modest but honest design-probe paper: the layered chat affordances are the contribution, and the evaluation is formative—treat it as a design exploration, not a validated system claim. read the letter →

arxiv 2506.03052 v1 pith:JQDHFSH5 submitted 2025-06-03 cs.HC

classification cs.HC
keywords conversationaluserinterfacedesignfeedbackprinciplessharedrepresentationscaffoldingreflectionnovicedesignersresearchthrough
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Feedstack, a conversational interface that layers bookmarks, chapters, and highlights onto an ordinary design-feedback chat with an AI. Its central claim is that these layered structures form a shared representation of the conversation that makes implicit design principles explicit, helping novice designers notice and reflect on what the feedback is really about. Two formative studies, each with eight beginners, produced reports of participants discovering principles they had overlooked and revisiting the conversation after it ended to apply the feedback. The paper frames this as an early design probe, not a conclusive evaluation.

What carries the argument

The load-bearing mechanism is the set of layered affordances wrapped around the chat transcript. Bookmarks place visible markers on the conversation's scrub bar wherever a design principle is discussed, so users can jump between instances. Chapters are LLM-generated accordions, each tied to a design principle, containing a definition, an explanation of how the principle applies to the user's design, and key terms; their opacity increases as the principle is discussed more. Highlights color key terms in the chat, and principle toggles switch them on and off. Together these layers are the system's shared representation of the conversation, giving the user and the AI common ground.

What would settle it

Have design experts annotate a set of real feedback conversations for the design principle each turn discusses, then run Feedstack's real-time chapter generator over the same conversations; if the generated chapters systematically miss or mislabel principles, the shared representation would mislead users rather than scaffold them.

Watch

Extended reading notes

Core claim

Feedstack's central claim is that a feedback conversation does not have to live only as a linear transcript. By bookmarking where each design principle is discussed, generating chapter summaries of those principles as the conversation unfolds, and highlighting key terms in place, the interface externalizes the tacit structure of a critique. Users can jump from a bookmark to an earlier 'Balance' moment, watch the Balance chapter expand beside the chat, and see how frequently each principle has come up through the chapter's opacity. That externalization is what lets a novice zoom out from turn-by-turn dialogue and grasp the larger themes, including principles that have not yet been discussed.

Load-bearing premise

The scaffolding effect rests on the AI reliably recognizing when a design principle is being discussed and generating accurate chapter content, a step the paper never tested with a live, automatically generated conversation.

Editorial extensions

If this is right

  • Novice designers can become aware of principles that were never named outright in the feedback, such as consistency or accessibility.
  • After the chat ends, the transcript becomes a navigable reference, so users can return to specific feedback instances while revising their designs.
  • Seeing a principle discussed in multiple bookmarked places supports comparing how the same principle plays out in different parts of a design.
  • Opacity and suggested emerging topics can point users toward design principles the conversation has so far neglected.
  • These benefits can appear without resorting to a rigid, wizard-driven dialogue, preserving the open-ended feel of a real critique.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The step most worth stress-testing is the real-time chapter generator: if the language model mislabels which principle a turn discusses, the whole shared representation would teach the wrong structure to learners.
  • The same layered-representation idea could transfer to other apprenticeship conversations, such as code review or scientific peer review, where the underlying principles are tacit for novices.
  • Opacity could be reused as a live coverage dashboard, making it easy to test whether a conversation has become unbalanced toward a few principles.
  • A natural follow-up experiment would measure learning, not just preference: give two groups of novices the same feedback, with and without the layered panel, and compare how accurately they can recall and apply the principles.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces Feedstack, a speculative conversational user interface that layers structured representations—bookmarks, chapters, highlights, principle toggles—on top of an unstructured chatbot feedback conversation. The authors report two formative user studies (n=8 each) with novice designers, the first observing interactions with ChatGPT and the second a think-aloud study with a pre-populated Feedstack conversation. The central claim is that these layered structures act as a shared representation between user and AI, surfacing implicit design principles, supporting exploration and reflection, and helping novices connect feedback instances to broader design principles. The paper is explicitly positioned as a research-through-design artifact and design probe rather than a controlled evaluation.

Significance. If the central claim holds, Feedstack contributes a concrete interaction pattern for educational CUIs: adding non-linear structure around a linear chat without destroying conversational flow. The paper's strengths are its grounding in established design literature (variation theory, dual-coding theory, shared representations, information scent), its clear and honest framing as a design probe, its two complementary small studies, and its use of participant quotes to illustrate design insights. The prototype is functional and the design rationale is reproducible. However, the significance is limited by the lack of a live-system evaluation and by the absence of any comparison condition, so the contribution is best understood as a design exploration that generates hypotheses rather than as evidence that the real-time pipeline benefits learning.

major comments (3)
  1. [4.2, 5.1, 6.3] The central claim that the layered structures constitute a shared representation that 'surfaces user intent and reveals underlying design principles' depends on the real-time LLM chapter-generation pipeline described in Section 4.2, but the user study in Section 5.1 'populated the conversation as text in the prototype' rather than exercising live generation. Consequently, none of the participant quotes in Section 5.2, such as P6 noticing Consistency or P7 noticing Balance instances, provide evidence about the accuracy, timing, or helpfulness of automatically generated chapters, learning materials, or opacity updates. If the LLM misassigns principles or generates plausible but incorrect definitions, the scaffolding could mislead novices rather than support them. The authors should either test the live pipeline with real conversations or explicitly limit the paper's contribution to the interaction design of pre-structured content, and temper the abstract and conclusion claims accordingly.
  2. [5.1, 5.2, 6.1.1] The study is an eight-participant think-aloud session without a comparison condition, and all reported benefits are self-reported perceptions rather than observed learning or changes in design behavior. For example, P2's 'teach them to fish' quote and P8's statement about future use are intentions, not evidence of improved understanding. The paper already frames itself as research-through-design, which is acceptable for generating hypotheses, but the conclusion that the tool 'helps novice designers become more aware' (Section 6.1.1) goes beyond what this design can support. I recommend replacing knowledge-claim language with design-insight language, e.g., 'participants reported that the affordances directed their attention to previously unnoticed principles.'
  3. [4.2, 5.1] The five-principle taxonomy (Accessibility, Consistency, Contrast, Balance, Alignment and Spacing) is presented as a fixed, sufficient set for the target feedback conversations, but no analysis is provided to show that this taxonomy covers the actual feedback utterances in the two studies. Since chapter opacity, bookmarks, and highlights are defined relative to these principles, an incomplete or mismatched taxonomy would systematically hide relevant feedback topics. The paper should report how the taxonomy was derived and whether any utterances in the formative study fell outside it, or acknowledge this as a limitation and an avenue for future work.
minor comments (5)
  1. [1] In the introduction, the sentence 'These ‘shared representations also serve to scaffold...' has an unclosed quotation mark and should read 'These shared representations also serve...'.
  2. [2.1] The phrase 'Many students now turning to general purpose chatbots' is missing an auxiliary verb and should read 'Many students are now turning...'.
  3. [2.2] The phrase 'conversations;for example' is missing a space after the semicolon; it should read 'conversations; for example'.
  4. [Figure 1 caption] The caption labels panels A-F, and the body text refers to 'the design panel A'; considering the chat panel is also a design element, renaming A to 'artifact panel' would reduce potential confusion.
  5. [6.2] Figure 3 is introduced after the discussion of Emerging Topics, Conversational Cues, and Referencing Conversation, but the text does not explicitly link the labeled elements L and M in the figure to these feature names; a brief cross-reference in the text would improve clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the design rationale draws on external prior work and the two studies are distinct; no fitted parameter or self-cited theorem is presented as a prediction.

full rationale

This paper makes no formal derivation or quantitative prediction. Feedstack's design goals are motivated by the authors' formative ChatGPT study and by external literature (e.g., Heer's shared representations, variation theory, dual-coding theory, and prior structured-feedback systems such as Decipher and Voyant). The second user study evaluates the prototype with a pre-populated conversation, so participants experienced hand-constructed content rather than the live LLM chapter-generation pipeline; this is a real validity limitation that the authors acknowledge in Section 6.3 ('the goal is not to validate hypotheses or to offer conclusive evidence'), but it is a missing evaluation, not a circular reduction. The authors' self-citations (e.g., [16,17,18,37,38]) are used as background claims about student chatbot use and prior feedback tools, not as the sole warrant for Feedstack's effectiveness. Consequently, there is no fitted input renamed as a prediction, no self-citation chain that forces the conclusion, and no definitional equivalence between the claimed outcome and the inputs. The central claim remains an exploratory, design-probe hypothesis with independent content.

Assumptions & free parameters 0 free parameters · 4 assumptions · 3 invented entities

No numerical free parameters are present. The claims rest on assumptions drawn from prior work on feedback and learning, plus two unverified design commitments: the fixed five-principle taxonomy and the real-time LLM chapter extraction. Future features are listed as invented entities because they are proposed without validation data.

assumptions (4)
  • domain assumption Structured feedback improves comprehension and application for novice learners.
    The design goals in Section 3 depend on this prior-work claim; the paper does not independently re-establish it.
  • ad hoc to paper The fixed five-principle taxonomy (Accessibility, Consistency, Contrast, Balance, Alignment and Spacing) is sufficient for the target feedback conversations.
    Section 4.2 makes this a design commitment based on [23]; it constrains what the system can surface.
  • domain assumption An LLM can reliably detect design principles in conversation and generate accurate chapter content in real time.
    Section 4.2 relies on this for chapter generation, but the second study used a pre-populated conversation, so extraction accuracy was not tested.
  • domain assumption Adding peripheral panels and visual layers does not unacceptably disrupt the conversational experience.
    This is the core premise behind D1 in Section 3; the studies did not measure conversational flow against a linear-chat baseline.
invented entities (3)
  • Emerging Topics (opacity extension)
    purpose: Surface related but undiscussed design principles as information scent.
    Proposed in Section 6.2.1; no user-study evidence is presented.
  • Conversational Cues
    purpose: Suggest conversational turns to guide users toward new topics or clarify feedback.
    Proposed in Section 6.2.2; not evaluated.
  • Referencing Conversation
    purpose: Provide a navigation area in Chapters that lists feedback instances for each principle.
    Described in Section 6.2.3 and shown in Figure 3, but not covered by the reported studies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Feedstack: Layering Structured Representations over Unstructured Feedback to Scaffold Human AI Conversation." pith.science (2026). https://pith.science/paper/JQDHFSH5

@misc{pith2026250603052,
  author       = {Pith},
  title        = {Pith review of: Feedstack: Layering Structured Representations over Unstructured Feedback to Scaffold Human AI Conversation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JQDHFSH5}},
  note         = {Machine review of arXiv:2506.03052}
}
read the original abstract

Many conversational user interfaces facilitate linear conversations with turn-based dialogue, similar to face-to-face conversations between people. However, digital conversations can afford more than simple back-and-forth; they can be layered with interaction techniques and structured representations that scaffold exploration, reflection, and shared understanding between users and AI systems. We introduce Feedstack, a speculative interface that augments feedback conversations with layered affordances for organizing, navigating, and externalizing feedback. These layered structures serve as a shared representation of the conversation that can surface user intent and reveal underlying design principles. This work represents an early exploration of this vision using a research-through-design approach. We describe system features and design rationale, and present insights from two formative (n=8, n=8) studies to examine how novice designers engage with these layered supports. Rather than presenting a conclusive evaluation, we reflect on Feedstack as a design probe that opens up new directions for conversational feedback systems.

Figures

Figures reproduced from arXiv: 2506.03052 by the authors.

Figure 1
Figure 1. The design panel A displays the user’s design artifact. The chat panel B contains the chat where the user can interact with the LLM-powered design expert to receive and explore feedback. The chapters panel C contains the interactive chapters D. On the top of the chat panel is the principle toggles E and on the left of the chat panel is the bookmarks F. D2. Support Exploration Enable users to connect individual piece… view at source ↗
Figure 2
Figure 2. The chat initially anchors at the bottom of the panel [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. The resulting web-based Feedstack system. It incorporates new features informed by recent findings. These include [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

39 extracted references · 31 canonical work pages

  1. [1]

    Salvatore Andolina, Hendrik Schneider, Joel Chan, Khalil Klouche, Giulio Jacucci, and Steven Dow. 2017. Crowdboard: augmenting in-person idea generation with real-time crowds. In Proceedings of the 2017 ACM SIGCHI Conference on Creativity and Cognition . 106–118

  2. [2]

    Alasdair Blair, Steven Curtis, Mark Goodwin, and Sam Shields. 2013. What Feedback do Students Want? Politics 33, 1 (Feb. 2013), 66–79. https://doi.org/10.1111/j.1467-9256.2012.01446.x

  3. [3]

    Kirsten Boehner, William Gaver, and Andy Boucher. 2012. Probes. In Inventive methods. Routledge, 185–201

  4. [4]

    Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative research in sport, exercise and health 11 (2019)

  5. [5]

    Joel Chan, Zijian Ding, Eesh Kamrah, and Mark Fuge. 2024. Formulating or Fixating: Effects of Examples on Problem Solving Vary as a Function of Example Presentation Interface Design. In Proceedings of the CHI Conference on Human Factors in Computing Systems . ACM, Honolulu HI USA, 1–16. https://doi.org/10.1145/3613904.3642653

  6. [6]

    Ed H Chi, Peter Pirolli, Kim Chen, and James Pitkow. 2001. Using information scent to model user information needs and actions and the Web. In Proceedings of the SIGCHI conference on Human factors in computing systems . 490–497

  7. [7]

    Jim Clark and Allan Paivio. 1991. Dual Coding Theory and Education. Educational Psychology Review 3 (09 1991), 149–210. https://doi.org/10.1007/ BF01320076 9 CUI ’25, July 8–10, 2025, Waterloo, ON, Canada Hannah Vy Nguyen et al

  8. [8]

    Eureka Foong, Darren Gergle, and Elizabeth M. Gerber. 2017. Novice and Expert Sensemaking of Crowdsourced Design Feedback. Proc. ACM Hum.-Comput. Interact. 1, CSCW, Article 45 (Dec. 2017), 18 pages. https://doi.org/10.1145/3134680

Show all 39 references
  1. [9]

    C Ailie Fraser, Tricia J Ngoon, Ariel S Weingarten, Mira Dontcheva, and Scott Klemmer. 2017. CritiqueKit: A mixed-initiative, real-time interface for improving feedback. In Adjunct Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology . 7–9

  2. [10]

    Marcel Gohsen, Johannes Kiesel, Mariam Korashi, Jan Ehlers, and Benno Stein. 2023. Guiding Oral Conversations: How to Nudge Users Towards Asking Questions?. In Proceedings of the 2023 Conference on Human Information Interaction and Retrieval (Austin, TX, USA) (CHIIR ’23). Asso...

  3. [11]

    Lingyuan Gu, Rongjin Hunag, and Ference Marton. 2004. Teaching with variation: A Chinese way of promoting effective mathematics learning. In How Chinese learn mathematics: Perspectives from insiders . World Scientific, 309–347

  4. [12]

    John Hattie and Helen Timperley. 2007. The Power of Feedback. Review of Educational Research 77, 1 (March 2007), 81–112

  5. [13]

    Jeffrey Heer. 2019. Agency plus automation: Designing artificial intelligence into interactive systems. Proceedings of the National Academy of Sciences 116, 6 (2019), 1844–1850

  6. [14]

    Pamela J Hinds. 1999. The curse of expertise: The effects of expertise and debiasing methods on prediction of novice performance. Journal of experimental psychology: applied 5, 2 (1999), 205

  7. [15]

    Hinds, Michael Patterson, and Jeffrey Pfeffer

    Pamela J. Hinds, Michael Patterson, and Jeffrey Pfeffer. 2001. Bothered by abstraction: The effect of expertise on knowledge transfer and subsequent novice performance. Journal of Applied Psychology 86, 6 (2001), 1232–1243. https://doi.org/10.1037/0021-9010.86.6.1232

  8. [16]

    Irene Hou, Sophia Mettille, Owen Man, Zhuo Li, Cynthia Zastudil, and Stephen MacNeil. 2024. The Effects of Generative AI on Computing Students’ Help-Seeking Preferences. In Proceedings of the 26th Australasian Computing Education Conference (Sydney, NSW, Australia)(ACE ’24). A...

  9. [17]

    Irene Hou, Hannah Vy Nguyen, Owen Man, and Stephen MacNeil. 2025. The Evolving Usage of GenAI by Computing Students. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 2 (SIGCSE TS 2025) . ACM, 1481–1482. https://doi.org/10.1145/3641555.3705266

  10. [18]

    Ziheng Huang, Kexin Quan, Joel Chan, and Stephen MacNeil. 2023. CausalMapper: Challenging designers to think in systems with Causal Maps and Large Language Model. In Proceedings of the 15th Conference on Creativity and Cognition . 325–329

  11. [19]

    Mohit Jain, Pratyush Kumar, Ramachandra Kota, and Shwetak N. Patel. 2018. Evaluating and Informing the Design of Chatbots. In Proceedings of the 2018 Designing Interactive Systems Conference (DIS ’18) . ACM, 12 pages. https://doi.org/10.1145/3196709.3196735

  12. [21]

    Markus Krause, Tom Garncarz, JiaoJiao Song, Elizabeth M Gerber, Brian P Bailey, and Steven P Dow. 2017. Critique style guide: Improving crowdsourced design feedback with a natural language model. In Proceedings of the 2017 CHI conference on human Factors in computing systems

  13. [22]

    Gyeong-Geon Lee and Xiaoming Zhai. 2024. Realizing Visual Question Answering for Education: GPT-4V as a Multimodal AI. ArXiv abs/2405.07163 (2024). https://api.semanticscholar.org/CorpusID:269757576

  14. [23]

    William Lidwell, Kritina Holden, and Jill Butler. 2020. Universal Principles of Design (3rd ed.). Rockport Publishers, Beverly, MA

  15. [24]

    Li Liu, Rama Subbareddy, and C. G. Raghavendra. 2022. AI Intelligence Chatbot to Improve Students Learning in the Higher Education Platform. Journal of Interconnection Networks 22 (2022). https://doi.org/10.1142/S0219265921430325

  16. [25]

    Kurt Luther, Amy Pavel, Wei Wu, Jari-lee Tolentino, Maneesh Agrawala, Björn Hartmann, and Steven P Dow. 2014. CrowdCrit: crowdsourcing and aggregating visual design critique. In Proceedings of the companion publication of the 17th ACM conference on Computer supported cooperati...

  17. [26]

    Kurt Luther, Jari-Lee Tolentino, Wei Wu, Amy Pavel, Brian P Bailey, Maneesh Agrawala, Björn Hartmann, and Steven P Dow. 2015. Structuring, aggregating, and evaluating crowdsourced design critique. In Proceedings of the 18th ACM conference on computer supported cooperative work...

  18. [27]

    Lauren Margulieux, Paul Denny, Kathryn Cunningham, Michael Deutsch, and Benjamin R Shapiro. 2021. When wrong is right: The instructional power of multiple conceptions. In Proceedings of the 17th ACM Conference on International Computing Education Research . 184–197

  19. [28]

    Chan, Antonio Garcia-Cabot, Eva Garcia-Lopez, and Héctor Amado-Salvatierra

    Mónica De La Roca, Miguel M. Chan, Antonio Garcia-Cabot, Eva Garcia-Lopez, and Héctor Amado-Salvatierra. 2024. The impact of a chatbot working as an assistant in a course for supporting student learning and engagement. Computer Applications in Engineering Education 32, 5 (2024...

  20. [29]

    Donald A Schön. 2017. The reflective practitioner: How professionals think in action . Routledge

  21. [30]

    Stefan Sonderegger and Sabine Seufert. 2022. Chatbot-mediated Learning: Conceptual Framework for the Design of Chatbot Use Cases in Education. In Proceedings of the 14th International Conference on Computer Supported Education - Volume 1: CSEDU . INSTICC, SciTePress, 207–215. ...

  22. [31]

    Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: Enabling multilevel exploration and sensemaking with large language models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–18

  23. [32]

    Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . ACM, San Francisco CA USA, 1–18. http...

  24. [33]

    Olia Tsivitanidou and Andri Ioannou. 2021. Envisioned Pedagogical Uses of Chatbots in Higher Education and Perceived Benefits and Challenges. In Learning and Collaboration Technologies: Games and Virtual Environments for Learning , Panayiotis Zaphiris and Andri Ioannou (Eds.)....

  25. [34]

    Helen Wauck, Yu-Chun Yen, Wai-Tat Fu, Elizabeth Gerber, Steven P Dow, and Brian P Bailey. 2017. From in the class or in the wild? Peers provide better design feedback than external crowds. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems . 5580–5591

  26. [35]

    Yangyu Xiao and Yuying Zhi. 2023. An Exploratory Study of EFL Learners’ Use of ChatGPT for Language Learning Tasks: Experience and Perceptions. Languages 8 (09 2023), 212. https://doi.org/10.3390/languages8030212

  27. [36]

    Anbang Xu, Shih-Wen Huang, and Brian Bailey. 2014. Voyant: generating structured feedback on visual designs using a crowd of non-experts. In Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing . ACM, Baltimore Maryland USA, 1433–144...

  28. [37]

    Dow, Elizabeth Gerber, and Brian P

    Yu-Chun Grace Yen, Steven P. Dow, Elizabeth Gerber, and Brian P. Bailey. 2017. Listen to Others, Listen to Yourself: Combining Feedback Review and Reflection to Improve Iterative Design. In Proceedings of the 2017 ACM SIGCHI Conference on Creativity and Cognition . ACM, Singap...

  29. [38]

    Yu-Chun Grace Yen, Joy O Kim, and Brian P Bailey. 2020. Decipher: an interactive visualization tool for interpreting unstructured design feedback from multiple providers. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13

  30. [39]

    Alvin Yuan, Kurt Luther, Markus Krause, Sophie Isabel Vennix, Steven P Dow, and Bjorn Hartmann. 2016. Almost an expert: The effects of rubrics and expertise on perceived value of crowdsourced design critiques. In Proceedings of the 19th ACM Conference on Computer-Supported Coo...

  31. [40]

    John Zimmerman, Jodi Forlizzi, and Shelley Evenson. 2007. Research through design as a method for interaction design research in HCI. In Proceedings of the SIGCHI Conference on Human factors in Computing Systems . 493–502. 11

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.