REVIEW 3 major objections 2 minor 1 references
Explainability-in-Action: Enabling Expressive Manipulation and Tacit Understanding by Bending Diffusion Models in ComfyUI
T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read By exposing a text-to-image diffusion model's internal stages as manipulable nodes in ComfyUI, this paper argues that artists can learn how each component shapes the output through hands-on 'model bending,' making tacit intuition a working
desk verdict Promising creative-XAI framing plus a concrete plugin, but the core claim about artists building intuition is asserted, not shown in the material we can read. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the plugin's mapping of diffusion-model internals onto ComfyUI's node graph, which turns each manipulable component into a visible stage in a live pipeline. This supports a reflection-in-action loop—acting on the model, observing the output, and adjusting based on what is seen—so that each intervention is also a probe that reveals the component's effect. The node graph is not just an interface convenience; it is what makes the model's structure available as hands-on material.
What would settle it
A controlled study in which artists with no prior knowledge of the plugin manipulate a single component and then predict the resulting image change; if their predictions are no better than chance after a fixed hands-on period, the claimed intuition development does not occur.
Extended reading notes
Core claim
The paper's central claim is that a large text-to-image diffusion model can be treated as a creative material—something an artist comes to understand through use—rather than as an opaque generator, provided its internal parts are exposed for direct manipulation. To make this concrete, the authors build a 'model-bending and inspection plugin' inside ComfyUI's node-based interface, where the model's stages are laid out as visible, connectable nodes that the artist can alter and immediately re-run. The intended effect is not a verbal explanation of what each component does, but a working, tacit intuition formed through repeated cycles of adjustment and observation. In this view, explainability
Load-bearing premise
The claim depends on the assumption that the changes an artist makes through the plugin reliably alter the output in ways that map to the component being manipulated, and that the demonstration by the plugin's designers transfers to other artists who do not share the designers' expertise.
Editorial extensions
If this is right
- If the claim holds, explainability for large generative models need not require peeling back the model to reveal its inner logic; exposing manipulable parts and letting the user experiment is enough to produce working understanding.
- Artists can exercise agency over large text-to-image systems by bending components rather than only prompting them, making the output a product of directed modification.
- The approach gives a practical criterion for explainable AI in creative tools: an explanation is useful if it increases the artist's ability to predict and control output changes.
- Long-term practice with such a plugin should accumulate into tacit knowledge that carries over to new models or interfaces structured the same way.
Reading between the lines
- If manipulation-based understanding is durable, it predicts that artists who train on this plugin will outperform explanation-reading peers when asked to produce a specified image change without the plugin's live feedback; the paper itself does not run this comparison.
- The same node-bending strategy could be applied to other generative architectures with composable internal stages, turning a question about model interpretability into a question about interface design.
- Instrumenting the plugin to log manipulation histories would offer a testable signature of intuition: skilled artists should make fewer, more targeted interventions than novices to reach the same output.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a craft-based approach to explainable AI for creative practice, arguing that large diffusion models can be treated as creative materials when their internal components are exposed and manipulable. It introduces a 'model-bending and inspection plugin' for ComfyUI, a node-based interface, and claims that by interactively manipulating different parts of a generative model, artists can develop an intuition about how each component influences the output. The abstract frames this as 'explainability-in-action' grounded in Schön's reflection-in-action. However, the supplied full text is encoding-corrupted, and the only readable evidence is the abstract; no user study, measures, or comparison baseline are reported.
Significance. If the behavioral claim were empirically established, the work would contribute a valuable alternative to transparency-centric XAI: it treats explanation as engagement with model internals in sustained practice, aligning with recent calls for material-centered AI. The plugin itself, if functional, could be a useful tool for artists. The paper also connects XAI to craft and tacit knowledge, which is a conceptual contribution. The paper's strengths are its coherent conceptual framing and concrete artifact proposal; it does not, in the readable portion, ship reproducible code, machine-checked proofs, or falsifiable predictions. The central demonstration, however, is unverified and currently rests on the authors' own expert engagement rather than on evidence from the target population of artists.
major comments (3)
- [Abstract, closing sentence] The abstract asserts: 'by interactively manipulating different parts of a generative model, artists can develop an intuition about how each component influences the output.' This is a behavioral claim about artists' tacit learning. The manuscript provides no user study, no structured observation, no pre/post measure, and no comparison baseline. The demonstration is performed by the designer-authors, who already know the model internals; their ability to read effects does not establish legibility for non-expert artists. This is load-bearing: the contribution is framed as 'explainability-in-action,' and without evidence of intuition development, the paper reduces to a plugin description.
- [Full text (all body sections after the abstract)] The supplied manuscript body is encoding-corrupted (mojibake); the implementation details, figures, and any evaluation are unreadable. This prevents verification of the plugin's mechanism, the claimed 'model bending' operations, and whether any empirical data exist. Even allowing for submission-formatting errors, a journal submission must be fully readable; this alone makes the manuscript unpublishable in its current form and prevents the referee from assessing the central claim.
- [No identifiable evaluation section] In the readable abstract and headers, there is no mention of an evaluation, user study, or results section. If such a section exists in the corrupted text, it is inaccessible, so the central causal claim—that interactive manipulation builds artist intuition—remains unsupported. For a cs.HC contribution, this requires at least a qualitative study with representative users or a systematically analyzed set of artist sessions. The authors must either add such evidence or explicitly reframe the paper as a systems/design proposal without the behavioral claim.
minor comments (2)
- [Abstract] The verb 'demonstrate' is used twice in the abstract, but the paper does not currently demonstrate the claimed outcome in an empirically verifiable way. Please specify whether this is a technical demonstration (the plugin runs) or an empirical demonstration (artists' intuition was measured).
- [Header/formatting] The body text is heavily garbled; please ensure the resubmission uses a clean, machine-readable PDF or text encoding. Figure and table callouts are unusable in the current version and must be restored.
Circularity Check
No circularity: no fitted parameters, equations, or self-citation chains; the unsupported intuition claim is an evidence gap, not a circular reduction.
full rationale
The paper is a qualitative HCI contribution with no quantitative derivation. The central assertion—'by interactively manipulating different parts of a generative model, artists can develop an intuition about how each component influences the output'—is an empirical claim, not the output of a model whose inputs already entail it. The plugin demonstration is the authors' own, and no user study or independent measure is reported, which is a correctness/validation weakness, not a definitional circularity. No fitted input is relabeled as a prediction, no uniqueness theorem is imported from previous work, and no self-citation bears the load. The paper should be evaluated for evidence sufficiency, not circularity; score 0.
Assumptions & free parameters
assumptions (2)
- domain assumption Schön's reflection-in-action is a valid epistemic model for how artists gain tacit understanding through manipulation
- domain assumption Diffusion model internals exposed by the plugin can be manipulated without breaking generation functionality
Cite this review
Pith. "Pith review of Explainability-in-Action: Enabling Expressive Manipulation and Tacit Understanding by Bending Diffusion Models in ComfyUI." pith.science (2026). https://pith.science/paper/K2NXAW2R
@misc{pith2026250807183,
author = {Pith},
title = {Pith review of: Explainability-in-Action: Enabling Expressive Manipulation and Tacit Understanding by Bending Diffusion Models in ComfyUI},
year = {2026},
howpublished = {\url{https://pith.science/paper/K2NXAW2R}},
note = {Machine review of arXiv:2508.07183}
}
read the original abstract
Explainable AI (XAI) in creative contexts can go beyond transparency to support artistic engagement, modifiability, and sustained practice. While curated datasets and training human-scale models can offer artists greater agency and control, large-scale generative models like text-to-image diffusion systems often obscure these possibilities. We suggest that even large models can be treated as creative materials if their internal structure is exposed and manipulable. We propose a craft-based approach to explainability rooted in long-term, hands-on engagement akin to Sch\"on's "reflection-in-action" and demonstrate its application through a model-bending and inspection plugin integrated into the node-based interface of ComfyUI. We demonstrate that by interactively manipulating different parts of a generative model, artists can develop an intuition about how each component influences the output.
Reference graph
Works this paper leans on
-
[1]
������������������������� �������� ���������� ������������ ��� ����� ������������� �� ������� �������� ������ �� ������� ����� �� ��������� ������ �� ����������� ���� ��� ���������� ����� ������ ���������� ������� ������ ��������������� �������� �������� ������ �� ����������� ���� ��� ���������� ����� ������ ���������� ������� ������ ��������������� �����...
arXiv 2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.