Pith. sign in

REVIEW 4 major objections 6 minor 33 references

Comparing Human and AI Performance in Visual Storytelling through Creation of Comic Strips: A Case Study

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that, given identical verbal instructions to recreate a three-panel comic strip, humans preserve the story while AI image generators produce polished but narratively incoherent images.

desk verdict Modest, honest case study that shows 2024 AI losing the narrative thread in a three-panel strip; the abstract overreaches, but the paper is worth a referee's time. read the letter →

arxiv 2507.18641 v1 pith:VA2V4D6U submitted 2025-05-27 cs.HC cs.CY

classification cs.HCcs.CY
keywords visualstorytellingcomicstripshuman-AIcomparisonAIimagegenerationnarrativecoherenceprompt-baseddrawingcasestudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether commercial AI image generators can match humans at visual storytelling, not just at making attractive pictures. Its test is a single famous three-panel Nancy cartoon strip, described only in words and given to three students with basic art training and to several AI tools. The paper claims that the humans reliably turned the instructions into a coherent story, while the best AI outputs looked professional but lost the story across panels. The authors present this as evidence that current AI still needs human judgment for effective visual narrative. They also report that a newer ChatGPT released in March 2025 already produces much better strips from the same prompt, so the observed gap may be shrinking.

What carries the argument

The load-bearing object is the three-panel Nancy strip from the book How to Read Nancy, chosen for its visual economy: every line and placement carries story information, so recreating it from text tests whether an agent can infer sequence, causality, and hidden intent. The verbal prompt in the paper is the other half of the machinery: the same instructions, with character names replaced by X and Y to block prior knowledge, were given to both humans and AI systems. Running both groups through that shared instruction set, and comparing the resulting panels, is what carries the argument.

What would settle it

Generate many outputs from current commercial image models using the exact prompt from the paper, have independent raters blind to source mark whether Panel 3 shows the girl with a hidden hose and the boy smirking and approaching, and compare coherence scores with the three human strips; if AI outputs are rated story-coherent as often as human ones, the paper's central claim would be overturned.

Watch

Extended reading notes

Core claim

Stated on its own terms, the discovery is that when a task requires holding a narrative together across three panels, humans and AI drawing systems diverge: humans preserve intent and causal sequence, while AI preserves surface style. The paper grounds this in side-by-side outputs for one prompt, showing four AI results chosen as the best from a larger set and three human results. The human strips depict both characters with identifiable actions and a setup-and-payoff structure; the AI strips render individual figures competently but do not reliably show the girl's hidden hose or the boy's oblivious approach, which are the beats that make the story. The authors conclude that AI excels at mimicking professional art but falls short at crafting coherent visual stories.

Load-bearing premise

The conclusion depends on treating the four hand-picked AI outputs, noted in Section 2.3, and the authors' qualitative reading of narrative coherence as representative of AI performance for one prompt and one strip.

Editorial extensions

If this is right

  • If AI systems cannot reliably hold a three-beat story together from a detailed prompt, commercial image generators are not yet suitable for unsupervised comic or storyboard production.
  • The same protocol can be rerun on other well-known strips to test whether the result depends on this particular story and prompt.
  • The March 2025 example indicates the capability is changing quickly, so repeated runs of this prompt can serve as a simple longitudinal benchmark for machine visual storytelling.
  • Because the three human participants were non-experts, the bar AI must reach is not professional comic craftsmanship but basic narrative competence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not quantify how the four 'best' AI examples were chosen; if the unshown outputs were worse, the displayed results may flatter AI, and if selection favored visual polish, it may understate AI's narrative failures.
  • Narrative coherence is judged by the authors themselves, so a natural next step is blind rating by independent annotators to test whether the gap is in the images or in the reading.
  • The prompt omits layout, style, and panel-relationship constraints, and commercial models are sensitive to wording; richer prompts might move AI performance substantially.
  • If this Nancy strip becomes a repeated benchmark, one could track simple plot-retention metrics, such as whether the hose and the smirk appear in Panel 3, across successive model releases.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper reports a case study in which three human students and several commercial AI image-generation systems were given the same textual instruction to recreate a three-panel Ernie Bushmiller comic strip. The authors qualitatively compare the outputs and conclude that AI systems are good at mimicking professional drawing styles but fail to produce coherent visual narratives, whereas humans are highly adept at turning instructions into meaningful stories. A brief disclaimer reports that a newer ChatGPT version released in March 2025 produced significantly improved strips from the same prompt.

Significance. If the central claim were supported, the paper would be a useful data point in the ongoing discussion of AI creativity, human-AI complementarity, and visual storytelling. The choice of a canonical, well-documented Nancy strip and the effort to anonymize characters as X and Y are thoughtful design decisions that could inform future benchmarks. Credit is also due for the authors' explicit acknowledgment in Section 3.1 that newer AI models already perform better on this task. However, because the empirical evidence consists of three human strips and four post hoc selected AI outputs, evaluated through unblinded subjective reading with no rubric, the paper does not establish its stated conclusion. Its current value is as a suggestive qualitative observation or a template for a more rigorous study, not as a demonstrated finding about 'AI systems' or 'humans' in general.

major comments (4)
  1. [Section 2.3] The AI evaluation set is not characterized: the text says 'A large set of outputs was generated using AI, but we present only the four best examples here.' The selection criterion for 'best' is not defined, and no information is given about the total number of outputs, the distribution of quality, or whether the selection was made before or after assessing narrative coherence. Because visual quality and narrative coherence are exactly the dimensions later judged, presenting only four selected outputs risks a selection bias that makes the comparison invalid. The paper should report the full output set, the sampling rule (e.g., random or pre-registered), and the number of outputs per model, or restrict all conclusions to the four displayed examples.
  2. [Section 2.4] The analysis is an unblinded, subjective appraisal by the authors. No scoring rubric, no independent raters, no inter-rater reliability measure, and no quantitative indicator (e.g., presence of required story elements, sequential consistency, or panel-to-panel continuity) are provided. The statement that 'AI struggled with sequential contexts' and that 'human-generated comics consistently captured the prompt's narrative' is therefore not verifiable from the reported evidence. A qualitative case study can be valuable, but the strength of the conclusion requires at least a clear annotation protocol and multiple blinded evaluators.
  3. [Section 3.1] The paper's own disclaimer states that a March 2025 version of ChatGPT 'began generating significantly improved cartoon strips from the same prompts.' This directly contradicts the abstract's categorical claim that 'AI systems ... struggle to create coherent visual stories' and the conclusion's statement that AI 'fall[s] short' in this respect. The central claim is therefore time-bound and model-specific even by the authors' admission. The conclusions should be explicitly limited to the specific models, prompts, and dates tested, and the newer output should either be included in the analysis or the scope should be narrowed accordingly.
  4. [Sections 2.3 and 2.4] The evidence base consists of three human participants and four selected AI outputs from heterogeneous systems (one Leonardo.Ai output, one custom-trained 'realisticVision' output, and two Dall-E outputs via ChatGPT). Yet the paper generalizes to 'AI systems' and 'humans' as categories. There is no justification that these instances represent their respective categories, no saturation argument, no variation in prompts, and no statistical treatment. The sample is too small and non-random to support the broad comparative claim; the manuscript should be reframed as an exploratory case study with explicit limitations or supplemented with a much larger, more systematic data collection.
minor comments (6)
  1. [Page 1 header] The title contains a typographical error: 'P ERFORMANCE' should be 'PERFORMANCE.'
  2. [Figure 4 caption] The caption lists three subfigure labels (a), (b), and (c) for four displayed images, with (c) described as 'Two examples created by OpenAI's Dall-E using ChatGPT.' The caption should clearly map each of the four panels to its producing model and label each panel separately.
  3. [Table 1] The Panel 3 instruction ends with 'On the right side, The character Y is walking...' where the capital 'T' after the comma is inconsistent; also specify clearly what 'the right side' refers to within the panel.
  4. [Section 2.3] The text refers to 'various AI tools' and later names Leonardo.Ai, a custom-trained 'realisticVision' model, and Dall-E via ChatGPT, but it does not specify which ChatGPT version was used in 2024. This information is essential for reproducibility.
  5. [References] The reference to Darroch (2017) appears to be a news article about a political dispute and does not seem related to the cartoon competition controversy mentioned in the introduction; please verify the citation or replace it with the intended source.
  6. [Section 2.3] The model is described as 'realisticVision' of 'Stable Baselines.' This likely refers to 'Stable Diffusion' rather than the reinforcement learning library 'Stable Baselines'; please correct the terminology.

Circularity Check

0 steps flagged · score 1.0 of 10

No meaningful circularity: the central claim rests on qualitative image comparison, not on a fitted input, self-citation chain, or definitional identity.

full rationale

This paper is an empirical case study rather than a derivation, so there is no equation chain to reduce to its own inputs. The central comparison is self-contained: the same prompt (Table 1) is given to three human participants and several AI tools, and the authors report qualitative observations in Sections 2.3 and 2.4. The claim that AI 'struggles' to create coherent visual stories is not a fitted parameter or a definitional consequence; it is an interpretation of generated images. The paper even includes a March 2025 ChatGPT counterexample in Section 3.1 that qualifies the claim, which further indicates the conclusion is not forced by construction. Self-citations, such as Akleman and Celik (2020), Akleman (2021), Akleman et al. (2015), and Dede et al. (2024), appear as background motivation about subtle expressive cues, but they are not used as proof of the empirical result. The main methodological risks—cherry-picking 'the four best' AI outputs and relying on unblinded subjective evaluation—concern validity and generalizability, not circularity. If anything, selecting the four best AI outputs would bias against the conclusion that AI struggles, so the conclusion is not an artifact of that selection. The score of 1 reflects only the mild presence of self-citations in the background and does not indicate a circular derivation.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim relies on the choice of the specific comic strip, the representativeness of the three human participants, and the authors' subjective evaluation of narrative coherence. No free parameters or invented entities are involved.

assumptions (3)
  • domain assumption The Nancy strip by Bushmiller is a suitable benchmark for visual storytelling quality
    The paper relies on the analysis in Newgarden and Karasik's book to assert the strip's quality, but does not justify why success on this single strip generalizes to visual storytelling broadly.
  • domain assumption The three human participants are representative of human artistic ability
    The paper uses three students with 'basic artistic training' and no knowledge of the strip, but provides no information about their variability or skill.
  • domain assumption The authors' qualitative judgment of narrative coherence is reliable
    The analysis in Section 2.4 is based on the authors' subjective reading of the outputs, with no rubric, blind evaluation, or inter-rater reliability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Comparing Human and AI Performance in Visual Storytelling through Creation of Comic Strips: A Case Study." pith.science (2026). https://pith.science/paper/VA2V4D6U

@misc{pith2026250718641,
  author       = {Pith},
  title        = {Pith review of: Comparing Human and AI Performance in Visual Storytelling through Creation of Comic Strips: A Case Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VA2V4D6U}},
  note         = {Machine review of arXiv:2507.18641}
}
read the original abstract

This article presents a case study comparing the capabilities of humans and artificial intelligence (AI) for visual storytelling. We developed detailed instructions to recreate a three-panel Nancy cartoon strip by Ernie Bushmiller and provided them to both humans and AI systems. The human participants were 20-something students with basic artistic training but no experience or knowledge of this comic strip. The AI systems used were popular commercial models trained to draw and paint like artists, though their training sets may not necessarily include Bushmiller's work. Results showed that AI systems excel at mimicking professional art but struggle to create coherent visual stories. In contrast, humans proved highly adept at transforming instructions into meaningful visual narratives.

Figures

Figures reproduced from arXiv: 2507.18641 by the authors.

Figure 1
Figure 1. An example highlights the role of subtle facial expressions in storytelling. In the left panel, the children [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The position and orientation of small props like droplets can alter perceived expressions (Akleman and Celik, [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The original "How to Read Nancy" comic strip drawn by Ernie Bushmiller that the prompt is derived from. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The four "comic strips" created by AI using the prompt. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: The three comic strips created by 20-something students with basic artistic training but no experience or [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: This is an example of the output generated using the prompt in the table with the new version of ChatGPT [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

33 extracted references · 33 canonical work pages

  1. [1]

    Akleman, E. (2021). Computing through time: Privacy. Computer , 54(08):9--9

  2. [2]

    and Celik, H

    Akleman, E. and Celik, H. (2020). Droplets. Sequentials , 1(4):4--29

  3. [3]

    Akleman, E., Franchi, S., Kaleci, D., Mandell, L., Yamauchi, T., Akleman, D., et al. (2015). A theoretical framework to represent narrative structures for visual storytelling. proceedings of bridges 2015: mathematics, Music, art, architecture, culture

  4. [4]

    T., and Weld, D

    Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., and Weld, D. (2021). Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems , pages 1--16

  5. [5]

    H., Ragnhildstveit, A., Sprockett, S., Barr, N., Christensen, A., and Seli, P

    Bellaiche, L., Shahi, R., Turpin, M. H., Ragnhildstveit, A., Sprockett, S., Barr, N., Christensen, A., and Seli, P. (2023). Humans versus ai: whether and why we prefer human-created compared to ai-created artwork. Cognitive Research: Principles and Implications , 8(1):42

  6. [6]

    Blair, P. (1995). Cartoon Animation: The Collector's Series . Walter Foster Publishing

  7. [7]

    and Roggen, D

    Bulling, A. and Roggen, D. (2011). Recognition of visual memory recall processes using eye movement analysis. In Proceedings of the 13th international conference on Ubiquitous computing , pages 455--464

  8. [8]

    Celik, H. (2011). On cartoon drawing. Bizim Gazete

Show all 33 references
  1. [9]

    C., Amershi, S., and Kamar, E

    Chang, J. C., Amershi, S., and Kamar, E. (2017). Revolt: Collaborative crowdsourcing for labeling machine learning datasets. In Proceedings of the 2017 CHI conference on human factors in computing systems , pages 2334--2346

  2. [10]

    Cumhuriyet (2024). 14. international turhan selcuk cartoon competition results. Cumhuriyet Newspaper: https://www.cumhuriyet.com.tr/turkiye/14-uluslararasi-turhan-selcuk-karikatur-yarismasi-sonuclandi-birinci-2208058

  3. [11]

    Darroch, G. (2017). Netherlands 'will pay the price' for blocking turkish visit – erdoğan. https://www.theguardian.com/world/2017/mar/12/netherlands-will-pay-the-price-for-blocking-turkish-visit-erdogan

  4. [12]

    A., Akleman, E., and Sezgin, M

    Dede, E., Agilonu, K. A., Akleman, E., and Sezgin, M. (2024). On the power of subtle expressive cues in the perception of human affects. arXiv preprint arXiv:2401.18013

  5. [13]

    Eisner, W. (2008). Comics and sequential art: Principles and practices from the legendary cartoonist . WW Norton & Company

  6. [14]

    Ekman, P. (1999). Facial expressions. Handbook of cognition and emotion , 16(301):e320

  7. [15]

    and Keltner, D

    Ekman, P. and Keltner, D. (1997). Universal facial expressions of emotion. Segerstrale U, P. Molnar P, eds. Nonverbal communication: Where nature meets culture , pages 27--46

  8. [16]

    and Oster, H

    Ekman, P. and Oster, H. (1979). Facial expressions of emotion. Annual review of psychology , 30(1):527--554

  9. [17]

    R., and Yan, H

    Fan, X., Shahid, A. R., and Yan, H. (2022). Edge-aware motion based facial micro-expression generation with attention mechanism. Pattern Recognition Letters , 162:97--104

  10. [18]

    and Thomas, F

    Johnston, O. and Thomas, F. (1981). The illusion of life: Disney animation . Disney Editions New York

  11. [19]

    Kamar, E., Hacker, S., and Horvitz, E. (2012). Combining human and machine intelligence in large-scale crowdsourcing. In AAMAS , volume 12, pages 467--474

  12. [20]

    Kotbas, M. (2024). Yapay zekalı ortalık toz duman.. at İzi İt İzine karışmaya başladı. Kotbaş ArtColors Blogspot: https://kotbasartcolors.blogspot.com/2024/05/yapay-zekal-ortalk-toz-duman-at-izi-it.html

  13. [21]

    Li, Y., Wei, J., Liu, Y., Kauttonen, J., and Zhao, G. (2022). Deep learning for micro-expression recognition: A survey. IEEE Transactions on Affective Computing , 13(4):2028--2046

  14. [22]

    Liu, Y., Akleman, E., Chen, J., et al. (2012). Never-ending storytelling with discrete-time markov processes. Proceedings of Bridges , pages 85--92

  15. [23]

    McCloud, S. (2006). Making comics: Storytelling secrets of comics, manga and graphic novels . Kitchen sink press Northampton, MA

  16. [24]

    and Martin, M

    McCloud, S. and Martin, M. (1993). Understanding comics: The invisible art , volume 106. Kitchen sink press Northampton, MA

  17. [25]

    K., and Yap, M

    Merghani, W., Davison, A. K., and Yap, M. H. (2018). A review on facial micro-expressions analysis: Datasets, features and metrics

  18. [26]

    and Karasik, P

    Newgarden, M. and Karasik, P. (1988). How to read nancy. In he Best of Ernie Bushmiller's Nancy , pages 98--105. Henry Holt/Comicana

  19. [27]

    and Karasik, P

    Newgarden, M. and Karasik, P. (2017). How to Read Nancy . Fantagraphics Books

  20. [28]

    A., Bachorowski, J.-A., and Fern \'a ndez-Dols, J.-M

    Russell, J. A., Bachorowski, J.-A., and Fern \'a ndez-Dols, J.-M. (2003). Facial and vocal expressions of emotion. Annual review of psychology , 54(1):329--349

  21. [29]

    Sindhura, S. P. and Abdul, A. (2021). Virtues and shortcomings of artificial intelligence in graphic design arena

  22. [30]

    K., Choudhary, C., Kunal, and Barnwal, P

    Trivedi, A., Kaur, E. K., Choudhary, C., Kunal, and Barnwal, P. (2023). Should ai technologies replace the human jobs? In 2023 2nd International Conference for Innovation in Technology (INOCON) , pages 1--6

  23. [31]

    https://leonardo.ai/

    URL-1 (2024). https://leonardo.ai/. [Online; date retrieved 18.09.2024]

  24. [32]

    VandenBos, G. R. (2007). APA dictionary of psychology. American Psychological Association

  25. [33]

    Viazovetskyi, Y., Ivashkin, V., and Kashin, E. (2020). Stylegan2 distillation for feed-forward image manipulation. In Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.-M., editors, Computer Vision -- ECCV 2020 , pages 170--186, Cham. Springer International Publishing

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.