REVIEW 4 major objections 6 minor 33 references
Comparing Human and AI Performance in Visual Storytelling through Creation of Comic Strips: A Case Study
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that, given identical verbal instructions to recreate a three-panel comic strip, humans preserve the story while AI image generators produce polished but narratively incoherent images.
desk verdict Modest, honest case study that shows 2024 AI losing the narrative thread in a three-panel strip; the abstract overreaches, but the paper is worth a referee's time. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the three-panel Nancy strip from the book How to Read Nancy, chosen for its visual economy: every line and placement carries story information, so recreating it from text tests whether an agent can infer sequence, causality, and hidden intent. The verbal prompt in the paper is the other half of the machinery: the same instructions, with character names replaced by X and Y to block prior knowledge, were given to both humans and AI systems. Running both groups through that shared instruction set, and comparing the resulting panels, is what carries the argument.
What would settle it
Generate many outputs from current commercial image models using the exact prompt from the paper, have independent raters blind to source mark whether Panel 3 shows the girl with a hidden hose and the boy smirking and approaching, and compare coherence scores with the three human strips; if AI outputs are rated story-coherent as often as human ones, the paper's central claim would be overturned.
Extended reading notes
Core claim
Stated on its own terms, the discovery is that when a task requires holding a narrative together across three panels, humans and AI drawing systems diverge: humans preserve intent and causal sequence, while AI preserves surface style. The paper grounds this in side-by-side outputs for one prompt, showing four AI results chosen as the best from a larger set and three human results. The human strips depict both characters with identifiable actions and a setup-and-payoff structure; the AI strips render individual figures competently but do not reliably show the girl's hidden hose or the boy's oblivious approach, which are the beats that make the story. The authors conclude that AI excels at mimicking professional art but falls short at crafting coherent visual stories.
Load-bearing premise
The conclusion depends on treating the four hand-picked AI outputs, noted in Section 2.3, and the authors' qualitative reading of narrative coherence as representative of AI performance for one prompt and one strip.
Editorial extensions
If this is right
- If AI systems cannot reliably hold a three-beat story together from a detailed prompt, commercial image generators are not yet suitable for unsupervised comic or storyboard production.
- The same protocol can be rerun on other well-known strips to test whether the result depends on this particular story and prompt.
- The March 2025 example indicates the capability is changing quickly, so repeated runs of this prompt can serve as a simple longitudinal benchmark for machine visual storytelling.
- Because the three human participants were non-experts, the bar AI must reach is not professional comic craftsmanship but basic narrative competence.
Reading between the lines
- The paper does not quantify how the four 'best' AI examples were chosen; if the unshown outputs were worse, the displayed results may flatter AI, and if selection favored visual polish, it may understate AI's narrative failures.
- Narrative coherence is judged by the authors themselves, so a natural next step is blind rating by independent annotators to test whether the gap is in the images or in the reading.
- The prompt omits layout, style, and panel-relationship constraints, and commercial models are sensitive to wording; richer prompts might move AI performance substantially.
- If this Nancy strip becomes a repeated benchmark, one could track simple plot-retention metrics, such as whether the hose and the smirk appear in Panel 3, across successive model releases.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a case study in which three human students and several commercial AI image-generation systems were given the same textual instruction to recreate a three-panel Ernie Bushmiller comic strip. The authors qualitatively compare the outputs and conclude that AI systems are good at mimicking professional drawing styles but fail to produce coherent visual narratives, whereas humans are highly adept at turning instructions into meaningful stories. A brief disclaimer reports that a newer ChatGPT version released in March 2025 produced significantly improved strips from the same prompt.
Significance. If the central claim were supported, the paper would be a useful data point in the ongoing discussion of AI creativity, human-AI complementarity, and visual storytelling. The choice of a canonical, well-documented Nancy strip and the effort to anonymize characters as X and Y are thoughtful design decisions that could inform future benchmarks. Credit is also due for the authors' explicit acknowledgment in Section 3.1 that newer AI models already perform better on this task. However, because the empirical evidence consists of three human strips and four post hoc selected AI outputs, evaluated through unblinded subjective reading with no rubric, the paper does not establish its stated conclusion. Its current value is as a suggestive qualitative observation or a template for a more rigorous study, not as a demonstrated finding about 'AI systems' or 'humans' in general.
major comments (4)
- [Section 2.3] The AI evaluation set is not characterized: the text says 'A large set of outputs was generated using AI, but we present only the four best examples here.' The selection criterion for 'best' is not defined, and no information is given about the total number of outputs, the distribution of quality, or whether the selection was made before or after assessing narrative coherence. Because visual quality and narrative coherence are exactly the dimensions later judged, presenting only four selected outputs risks a selection bias that makes the comparison invalid. The paper should report the full output set, the sampling rule (e.g., random or pre-registered), and the number of outputs per model, or restrict all conclusions to the four displayed examples.
- [Section 2.4] The analysis is an unblinded, subjective appraisal by the authors. No scoring rubric, no independent raters, no inter-rater reliability measure, and no quantitative indicator (e.g., presence of required story elements, sequential consistency, or panel-to-panel continuity) are provided. The statement that 'AI struggled with sequential contexts' and that 'human-generated comics consistently captured the prompt's narrative' is therefore not verifiable from the reported evidence. A qualitative case study can be valuable, but the strength of the conclusion requires at least a clear annotation protocol and multiple blinded evaluators.
- [Section 3.1] The paper's own disclaimer states that a March 2025 version of ChatGPT 'began generating significantly improved cartoon strips from the same prompts.' This directly contradicts the abstract's categorical claim that 'AI systems ... struggle to create coherent visual stories' and the conclusion's statement that AI 'fall[s] short' in this respect. The central claim is therefore time-bound and model-specific even by the authors' admission. The conclusions should be explicitly limited to the specific models, prompts, and dates tested, and the newer output should either be included in the analysis or the scope should be narrowed accordingly.
- [Sections 2.3 and 2.4] The evidence base consists of three human participants and four selected AI outputs from heterogeneous systems (one Leonardo.Ai output, one custom-trained 'realisticVision' output, and two Dall-E outputs via ChatGPT). Yet the paper generalizes to 'AI systems' and 'humans' as categories. There is no justification that these instances represent their respective categories, no saturation argument, no variation in prompts, and no statistical treatment. The sample is too small and non-random to support the broad comparative claim; the manuscript should be reframed as an exploratory case study with explicit limitations or supplemented with a much larger, more systematic data collection.
minor comments (6)
- [Page 1 header] The title contains a typographical error: 'P ERFORMANCE' should be 'PERFORMANCE.'
- [Figure 4 caption] The caption lists three subfigure labels (a), (b), and (c) for four displayed images, with (c) described as 'Two examples created by OpenAI's Dall-E using ChatGPT.' The caption should clearly map each of the four panels to its producing model and label each panel separately.
- [Table 1] The Panel 3 instruction ends with 'On the right side, The character Y is walking...' where the capital 'T' after the comma is inconsistent; also specify clearly what 'the right side' refers to within the panel.
- [Section 2.3] The text refers to 'various AI tools' and later names Leonardo.Ai, a custom-trained 'realisticVision' model, and Dall-E via ChatGPT, but it does not specify which ChatGPT version was used in 2024. This information is essential for reproducibility.
- [References] The reference to Darroch (2017) appears to be a news article about a political dispute and does not seem related to the cartoon competition controversy mentioned in the introduction; please verify the citation or replace it with the intended source.
- [Section 2.3] The model is described as 'realisticVision' of 'Stable Baselines.' This likely refers to 'Stable Diffusion' rather than the reinforcement learning library 'Stable Baselines'; please correct the terminology.
Circularity Check
No meaningful circularity: the central claim rests on qualitative image comparison, not on a fitted input, self-citation chain, or definitional identity.
full rationale
This paper is an empirical case study rather than a derivation, so there is no equation chain to reduce to its own inputs. The central comparison is self-contained: the same prompt (Table 1) is given to three human participants and several AI tools, and the authors report qualitative observations in Sections 2.3 and 2.4. The claim that AI 'struggles' to create coherent visual stories is not a fitted parameter or a definitional consequence; it is an interpretation of generated images. The paper even includes a March 2025 ChatGPT counterexample in Section 3.1 that qualifies the claim, which further indicates the conclusion is not forced by construction. Self-citations, such as Akleman and Celik (2020), Akleman (2021), Akleman et al. (2015), and Dede et al. (2024), appear as background motivation about subtle expressive cues, but they are not used as proof of the empirical result. The main methodological risks—cherry-picking 'the four best' AI outputs and relying on unblinded subjective evaluation—concern validity and generalizability, not circularity. If anything, selecting the four best AI outputs would bias against the conclusion that AI struggles, so the conclusion is not an artifact of that selection. The score of 1 reflects only the mild presence of self-citations in the background and does not indicate a circular derivation.
Assumptions & free parameters
assumptions (3)
- domain assumption The Nancy strip by Bushmiller is a suitable benchmark for visual storytelling quality
- domain assumption The three human participants are representative of human artistic ability
- domain assumption The authors' qualitative judgment of narrative coherence is reliable
Cite this review
Pith. "Pith review of Comparing Human and AI Performance in Visual Storytelling through Creation of Comic Strips: A Case Study." pith.science (2026). https://pith.science/paper/VA2V4D6U
@misc{pith2026250718641,
author = {Pith},
title = {Pith review of: Comparing Human and AI Performance in Visual Storytelling through Creation of Comic Strips: A Case Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/VA2V4D6U}},
note = {Machine review of arXiv:2507.18641}
}
read the original abstract
This article presents a case study comparing the capabilities of humans and artificial intelligence (AI) for visual storytelling. We developed detailed instructions to recreate a three-panel Nancy cartoon strip by Ernie Bushmiller and provided them to both humans and AI systems. The human participants were 20-something students with basic artistic training but no experience or knowledge of this comic strip. The AI systems used were popular commercial models trained to draw and paint like artists, though their training sets may not necessarily include Bushmiller's work. Results showed that AI systems excel at mimicking professional art but struggle to create coherent visual stories. In contrast, humans proved highly adept at transforming instructions into meaningful visual narratives.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Akleman, E. (2021). Computing through time: Privacy. Computer , 54(08):9--9
work page 2021
- [2]
-
[3]
Akleman, E., Franchi, S., Kaleci, D., Mandell, L., Yamauchi, T., Akleman, D., et al. (2015). A theoretical framework to represent narrative structures for visual storytelling. proceedings of bridges 2015: mathematics, Music, art, architecture, culture
work page 2015
-
[4]
Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., and Weld, D. (2021). Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems , pages 1--16
work page 2021
-
[5]
H., Ragnhildstveit, A., Sprockett, S., Barr, N., Christensen, A., and Seli, P
Bellaiche, L., Shahi, R., Turpin, M. H., Ragnhildstveit, A., Sprockett, S., Barr, N., Christensen, A., and Seli, P. (2023). Humans versus ai: whether and why we prefer human-created compared to ai-created artwork. Cognitive Research: Principles and Implications , 8(1):42
work page 2023
-
[6]
Blair, P. (1995). Cartoon Animation: The Collector's Series . Walter Foster Publishing
work page 1995
-
[7]
Bulling, A. and Roggen, D. (2011). Recognition of visual memory recall processes using eye movement analysis. In Proceedings of the 13th international conference on Ubiquitous computing , pages 455--464
work page 2011
-
[8]
Celik, H. (2011). On cartoon drawing. Bizim Gazete
work page 2011
Show all 33 references
-
[9]
C., Amershi, S., and Kamar, E
Chang, J. C., Amershi, S., and Kamar, E. (2017). Revolt: Collaborative crowdsourcing for labeling machine learning datasets. In Proceedings of the 2017 CHI conference on human factors in computing systems , pages 2334--2346
2017
-
[10]
Cumhuriyet (2024). 14. international turhan selcuk cartoon competition results. Cumhuriyet Newspaper: https://www.cumhuriyet.com.tr/turkiye/14-uluslararasi-turhan-selcuk-karikatur-yarismasi-sonuclandi-birinci-2208058
2024
-
[11]
Darroch, G. (2017). Netherlands 'will pay the price' for blocking turkish visit – erdoğan. https://www.theguardian.com/world/2017/mar/12/netherlands-will-pay-the-price-for-blocking-turkish-visit-erdogan
2017
-
[12]
A., Akleman, E., and Sezgin, M
Dede, E., Agilonu, K. A., Akleman, E., and Sezgin, M. (2024). On the power of subtle expressive cues in the perception of human affects. arXiv preprint arXiv:2401.18013
2024 arXiv
-
[13]
Eisner, W. (2008). Comics and sequential art: Principles and practices from the legendary cartoonist . WW Norton & Company
2008
-
[14]
Ekman, P. (1999). Facial expressions. Handbook of cognition and emotion , 16(301):e320
1999
-
[15]
and Keltner, D
Ekman, P. and Keltner, D. (1997). Universal facial expressions of emotion. Segerstrale U, P. Molnar P, eds. Nonverbal communication: Where nature meets culture , pages 27--46
1997
-
[16]
and Oster, H
Ekman, P. and Oster, H. (1979). Facial expressions of emotion. Annual review of psychology , 30(1):527--554
1979
-
[17]
R., and Yan, H
Fan, X., Shahid, A. R., and Yan, H. (2022). Edge-aware motion based facial micro-expression generation with attention mechanism. Pattern Recognition Letters , 162:97--104
2022
-
[18]
and Thomas, F
Johnston, O. and Thomas, F. (1981). The illusion of life: Disney animation . Disney Editions New York
1981
-
[19]
Kamar, E., Hacker, S., and Horvitz, E. (2012). Combining human and machine intelligence in large-scale crowdsourcing. In AAMAS , volume 12, pages 467--474
2012
-
[20]
Kotbas, M. (2024). Yapay zekalı ortalık toz duman.. at İzi İt İzine karışmaya başladı. Kotbaş ArtColors Blogspot: https://kotbasartcolors.blogspot.com/2024/05/yapay-zekal-ortalk-toz-duman-at-izi-it.html
2024
-
[21]
Li, Y., Wei, J., Liu, Y., Kauttonen, J., and Zhao, G. (2022). Deep learning for micro-expression recognition: A survey. IEEE Transactions on Affective Computing , 13(4):2028--2046
2022
-
[22]
Liu, Y., Akleman, E., Chen, J., et al. (2012). Never-ending storytelling with discrete-time markov processes. Proceedings of Bridges , pages 85--92
2012
-
[23]
McCloud, S. (2006). Making comics: Storytelling secrets of comics, manga and graphic novels . Kitchen sink press Northampton, MA
2006
-
[24]
and Martin, M
McCloud, S. and Martin, M. (1993). Understanding comics: The invisible art , volume 106. Kitchen sink press Northampton, MA
1993
-
[25]
K., and Yap, M
Merghani, W., Davison, A. K., and Yap, M. H. (2018). A review on facial micro-expressions analysis: Datasets, features and metrics
2018
-
[26]
and Karasik, P
Newgarden, M. and Karasik, P. (1988). How to read nancy. In he Best of Ernie Bushmiller's Nancy , pages 98--105. Henry Holt/Comicana
1988
-
[27]
and Karasik, P
Newgarden, M. and Karasik, P. (2017). How to Read Nancy . Fantagraphics Books
2017
-
[28]
A., Bachorowski, J.-A., and Fern \'a ndez-Dols, J.-M
Russell, J. A., Bachorowski, J.-A., and Fern \'a ndez-Dols, J.-M. (2003). Facial and vocal expressions of emotion. Annual review of psychology , 54(1):329--349
2003
-
[29]
Sindhura, S. P. and Abdul, A. (2021). Virtues and shortcomings of artificial intelligence in graphic design arena
2021
-
[30]
K., Choudhary, C., Kunal, and Barnwal, P
Trivedi, A., Kaur, E. K., Choudhary, C., Kunal, and Barnwal, P. (2023). Should ai technologies replace the human jobs? In 2023 2nd International Conference for Innovation in Technology (INOCON) , pages 1--6
2023
-
[31]
https://leonardo.ai/
URL-1 (2024). https://leonardo.ai/. [Online; date retrieved 18.09.2024]
2024
-
[32]
VandenBos, G. R. (2007). APA dictionary of psychology. American Psychological Association
2007
-
[33]
Viazovetskyi, Y., Ivashkin, V., and Kashin, E. (2020). Stylegan2 distillation for feed-forward image manipulation. In Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.-M., editors, Computer Vision -- ECCV 2020 , pages 170--186, Cham. Springer International Publishing
2020
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.