Pith. sign in

REVIEW 5 major objections 5 minor 75 references

Seeing the Forest and the Trees: Solving Visual Graph and Tree Based Data Structure Problems using Large Multimodal Models

T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper claims that off-the-shelf multimodal models, given only a rendered image and a simple zero-shot prompt, correctly solve most tree and many graph data-structure exam problems, led by GPT-4o at 87.6% on trees and Gemini 1.5 Flash…

desk verdict Useful benchmark and a credible 'LMMs already solve many visual tree problems' result, but exact accuracies need a parsing audit and error bars. read the letter →

arxiv 2412.11088 v1 pith:A6ACQZ7L submitted 2024-12-15 cs.AI cs.CVcs.CY

classification cs.AIcs.CVcs.CY
keywords largemultimodalmodelsvisualgraphproblemstreetraversalbenchmarkdatasetzero-shotpromptingcomputingeducationacademicintegrityassessment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Educators have leaned on image-based graph and tree problems as a way to keep AI assistants out of take-home and online exams. This paper argues that the defense has already failed: current large multimodal models, prompted with one short instruction and no examples, recover the correct traversal or adjacency list from a rendered image most of the time on trees and more than half the time on graphs. To make the measurement leakage-resistant, the authors generated 9,072 unique problems spanning binary trees, binary search trees, undirected graphs, and directed graphs, with controlled variation in node count, edge width, node color, and task type. GPT-4o reached 87.6% pass@3 on trees and Gemini 1.5 Flash 56.2% pass@3 on graphs. The paper concludes that visual presentation cannot be treated as a reliable assessment safeguard and that pedagogy must move toward forms of assessment that do not depend on rendering a problem in an image.

What carries the argument

The engine of the study is a programmatic benchmark generator. It produces 9,072 image-text-answer triples across four structure classes — binary tree, binary search tree, undirected graph, directed graph — with node counts from 3 to 9, two edge widths, two node colors, several node-value sets, and six operation prompts, with each expected answer computed from the structure used to render the image. This generator serves three roles: it removes the risk of training-data leakage, it creates controlled structural and aesthetic variation for the feature analysis, and it yields a reusable open-source tool for re-testing future LMMs. The evaluation machinery is zero-shot prompting plus pass@k scoring with regular-expression answer extraction, and a logistic-regression model over engineered graph and image features identifies which variations drive accuracy.

What would settle it

Human annotators should re-derive the expected traversal or adjacency list for a random sample of the benchmark images; if even a small share of images produces disagreement with the generator's ground truth, or if the regex extraction discards answers a human would call correct, the measured accuracies are not a clean measure of visual problem-solving.

Watch

Extended reading notes

Core claim

In the paper's own terms, the central finding is empirical: with a zero-shot user message that names the structure and asks for a Python-typed answer, the best tested LMMs can read a rendered diagram and produce the correct pre-order, in-order, post-order, breadth-first, or depth-first traversal, or the adjacency list, for a substantial fraction of programmatically generated, never-before-seen problems. GPT-4o achieves 77.8% pass@1 and 87.6% pass@3 on trees, and Gemini 1.5 Flash achieves 52.0% pass@1 and 56.2% pass@3 on graphs. Performance falls as the number of edges grows and is better on binary search trees than on directed graphs, which the authors attribute partly to the top-left-to-bottom-right patch reading of vision transformers. Because the prompting is deliberately minimal, the authors read these accuracies as a lower bound on what a student could obtain by iterating on prompts, making the result a statement about assessment integrity rather than only about model capability.

Load-bearing premise

The whole measurement assumes every rendered diagram is an unambiguous picture of the graph or tree whose traversal is the ground truth, and that a model answer that says the right thing in slightly different words is still counted as correct; if any rendering is ambiguous or any correct answer fails the extraction step, the reported percentages are biased.

Editorial extensions

If this is right

  • Visual graph and tree problems should no longer be treated as inherently AI-resistant assessment items; a bare prompt suffices to solve most tree questions at pass@3.
  • Because the reported numbers come from zero-shot prompting, they are a floor: students who iterate or use few-shot prompt engineering would plausibly do better, so the practical integrity risk exceeds the headline accuracies.
  • Accuracy is systematically sensitive to structural load: performance degrades as edge count, density, and degree histograms rise, so any future image-based assessment that wants to survive must push problems into high-complexity regimes.
  • The released generator and 9,072-problem benchmark give instructors and researchers a standardized way to re-measure new models as they appear, so the result is a moving baseline rather than a one-time snapshot.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: because no human verification of the rendered images or parse-failure rate is reported, the measurement's cleanliness is asserted rather than demonstrated; if valid model outputs are discarded by the extraction step, the true visual-reading rate would be higher than reported.
  • An untested implication is that the same problems posed as adjacency-matrix text rather than images would be nearly trivial for the same models; comparing those two conditions would isolate the visual-reading component of the errors.
  • The patch-order explanation suggests a concrete extension: rotating or mirroring the rendered trees and graphs should change accuracy in a predictable way if vision transformers parse images top-left to bottom-right, and the dataset could test that directly.
  • The authors stop at solving operations on given structures; an immediate next step is generating code from a diagram, which they flag as future work and which would widen the assessment threat to algorithm design questions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper introduces a benchmark of 9,072 programmatically generated graph and tree problems, evaluates eight large multimodal models (GPT-4o, GPT-4V, Gemini 1.5 Pro, Gemini 1.5 Flash, Gemini 1.0 Pro Vision, and Claude 3 Opus/Sonnet/Haiku) under zero-shot prompting, and reports pass@1 and pass@3 accuracies. The headline results are GPT-4o at 87.6% pass@3 on tree problems and Gemini 1.5 Flash at 56.2% pass@3 on graph problems. The authors also fit logistic regression models to engineered structural, aesthetic, and image features to answer questions about which variations influence accuracy, and they release the generator and dataset under an MIT license. The paper's central claim is that off-the-shelf LMMs can solve a substantial share of visually presented graph and tree exam questions, which has implications for assessment design in computing education.

Significance. If the reported measurements are valid, this is a useful and timely contribution to computing-education research. The open-source generator and benchmark address data-leakage concerns by creating novel items, and the systematic variation of node counts, edge widths, colors, layouts, and task types enables a more fine-grained analysis than prior visual Parsons-problem work. The paper explicitly positions image-based graph and tree questions as an assessment strategy and provides evidence that this strategy is already fragile. The inclusion of multiple model families and explicit zero-shot prompting is a practical strength. However, the significance of the headline accuracy numbers and the feature-importance findings depends on the validity of the rendering-to-ground-truth alignment, the answer extraction pipeline, and the statistical treatment of the point estimates, all of which are currently under-specified.

major comments (5)
  1. [Section 3.4, Table 1] The paper reports pass@1/pass@3 accuracies without reporting how many model responses failed the regular-expression extraction or what cleaning steps were applied. A correct answer that is not parseable is scored as wrong, while a parseable substring of an incorrect answer can be scored right, so the Table 1 point estimates are not a clean measurement of visual problem-solving until the parse-failure rate and a robustness check (e.g., alternative extraction or manual review of a sample) are reported per model.
  2. [Section 3.1.3, Figure 2] The ground-truth answers A_true are generated by the same program that renders the images, but the paper provides no independent check that the stored answer is entailed by the visible rendering. For tree traversals, in/pre/post-order depend on the root and the left/right child relation, which are properties of the internal tree, not necessarily of the matplotlib layout used to draw the image. Without a human spot-check of rendered images against A_true, or an audit of layout consistency, the headline tree accuracies (e.g., GPT-4o at 87.6% pass@3) may be biased by ambiguous or mismatched renderings.
  3. [Section 3.4, pass@3 definition] The definition of pass@3 is not specified: the paper does not state whether three independent samples are drawn per item at temperature 1.0, how per-item aggregation is computed, or how ties are broken. Furthermore, no confidence intervals or significance tests are reported, so interpretations of small differences in Table 1 are unsupported. A cluster-bootstrap confidence interval over items, or an equivalent, is needed for the comparative claims in Section 4.1.
  4. [Appendix C.1.6] The graph density formulas are printed incorrectly: directed density should be E/(V(V-1)) and undirected density should be 2E/(V(V-1)), but the manuscript writes edges in the denominator, making the ratio depend only on node count. If the implemented feature used these formulas, the density feature in the logistic regression is not density, and the RQ2 finding in Section 4.2.4 ('positive influence at lower and negative at higher densities') is not about density as defined. The formulas must be corrected and the feature analysis re-run or shown to be unaffected.
  5. [Section 5.1] The discussion claims that models achieve performance 'sufficient to pass traditional assessments,' but no human or student baseline and no passing threshold is provided. The 87.6% tree accuracy is hard to calibrate without knowing expected human performance on the same generated items; a small human study or a comparison to instructor-produced solutions is needed to support this pedagogical conclusion.
minor comments (5)
  1. [Section 3.1.3] The sentence 'The dataset is divided equally between graph (|D_G| = 4536) and tree (|D_G| = 4536)' uses the same symbol twice; the second should be |D_T| = 4536.
  2. [Section 4.4] The statement that 'Claude 3 Opus outperformed other models on directed graphs' contradicts Table 1, where Gemini 1.5 Flash and Gemini 1.5 Pro achieve much higher directed-graph accuracies; the sentence likely meant that Claude 3 Opus performed better on directed graphs than on undirected graphs.
  3. [Section 3.2] The sentence 'Preemptive sample evaluations on Gemini 1.0 Pro Vision show minimal differences in accuracy compared to the six prompts tested' does not specify what the six prompts were, the sample size used, or the magnitude of the differences; please include these details for reproducibility.
  4. [Section 3.3] Please provide exact model API identifiers and snapshot dates (e.g., gpt-4o-2024-05-13) in addition to the evaluation dates, since model versions can affect results.
  5. [ACM reference block] The ACM reference format line has placeholder text 'In ,.' and '14 pages'; this should be completed with the actual proceedings name and location before the camera-ready version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline accuracies are external measurements against programmatically generated ground truth, and the fitted feature-importance analysis is descriptive rather than definitional.

full rationale

The paper's central claim, that LMMs achieve the accuracies in Table 1 on visual graph and tree problems, is an empirical measurement rather than a derivation from fitted inputs. Ground-truth answers A_true are generated from the same networkx structures used to render the images, but model predictions A_model are produced independently by the models under zero-shot prompting, and accuracy is computed by comparing the two (Section 3.4). No parameter is fitted to the model outputs and then renamed as a prediction. The logistic-regression feature-importance analysis (Sections 3.4.1, 4.2, 4.3) is fitted to the evaluation results, but it is explicitly a post-hoc explanation of which features correlate with accuracy; it does not define the accuracy numbers and is not used to derive them. Self-citations such as [33] provide context from prior work on Parsons problems and are not load-bearing premises for the benchmark results. Concerns about regex parsing failures, layout ambiguity, and lack of human spot-checks are threats to measurement validity, not circularity, because they concern whether the external comparison is accurate rather than whether the conclusion is equivalent to its inputs by construction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The headline accuracy measurements depend on two front-end assumptions: that the rendered images faithfully present the recorded graph structure, and that regex parsing faithfully scores model outputs. Neither is directly verified with a parse-failure rate or a human spot-check. The feature-importance analysis adds fitted regression coefficients, but these are descriptive and not used to define the target accuracy values.

free parameters (3)
  • Sampling temperature = 1.0
    Set to 1.0, described as the most common default across models. This affects the stochastic pass@3 results but is not a parameter fitted to make the central claim true.
  • Logistic regression coefficients for feature importance
    Coefficients in the post-hoc logistic regression (Appendix D) are fitted to the evaluation results to rank feature influence for RQ2/RQ3; they are descriptive and are not used to derive the headline accuracy values.
  • PCA component count for image features
    Principal components retained from EfficientViT features are chosen based on variance explained; this affects the feature importance ranking but not the central accuracy measurements.
assumptions (4)
  • domain assumption Ground truth traversals and adjacency lists are computed correctly and unambiguously from the internal graph representation
    The generator (Section 3.1) builds images with networkx/matplotlib and records 'true' answers from the same data; if rendering introduces ambiguity, the ground truth may diverge from what a human or model sees.
  • domain assumption Regular-expression extraction faithfully captures model correctness
    Section 3.4 says answers are extracted with regular expressions; no parse-failure rate is reported, so the accuracy estimates may be inflated or deflated by formatting mismatches.
  • domain assumption Zero-shot prompting approximates realistic student usage
    Section 3.2 frames the zero-shot choice as the likely student behavior; the paper's pedagogy conclusions depend on this assumption, but it is not tested against actual student prompting behavior.
  • standard math Standard graph traversal definitions from CS curricula
    The tasks assume conventional definitions of in-order, pre-order, post-order, DFS, BFS, and adjacency lists from the ACM CS2013 curriculum reference used in Section 3.1.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Seeing the Forest and the Trees: Solving Visual Graph and Tree Based Data Structure Problems using Large Multimodal Models." pith.science (2026). https://pith.science/paper/A6ACQZ7L

@misc{pith2026241211088,
  author       = {Pith},
  title        = {Pith review of: Seeing the Forest and the Trees: Solving Visual Graph and Tree Based Data Structure Problems using Large Multimodal Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A6ACQZ7L}},
  note         = {Machine review of arXiv:2412.11088}
}
read the original abstract

Recent advancements in generative AI systems have raised concerns about academic integrity among educators. Beyond excelling at solving programming problems and text-based multiple-choice questions, recent research has also found that large multimodal models (LMMs) can solve Parsons problems based only on an image. However, such problems are still inherently text-based and rely on the capabilities of the models to convert the images of code blocks to their corresponding text. In this paper, we further investigate the capabilities of LMMs to solve graph and tree data structure problems based only on images. To achieve this, we computationally construct and evaluate a novel benchmark dataset comprising 9,072 samples of diverse graph and tree data structure tasks to assess the performance of the GPT-4o, GPT-4v, Gemini 1.5 Pro, Gemini 1.5 Flash, Gemini 1.0 Pro Vision, and Claude 3 model families. GPT-4o and Gemini 1.5 Flash performed best on trees and graphs respectively. GPT-4o achieved 87.6% accuracy on tree samples, while Gemini 1.5 Flash, achieved 56.2% accuracy on graph samples. Our findings highlight the influence of structural and visual variations on model performance. This research not only introduces an LMM benchmark to facilitate replication and further exploration but also underscores the potential of LMMs in solving complex computing problems, with important implications for pedagogy and assessment practices.

Figures

Figures reproduced from arXiv: 2412.11088 by the authors.

Figure 1
Figure 1. Overview of our benchmark dataset creation process, illustrating the transition from dataset construction to model [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The complete set of images from the dataset. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Zero-shot accuracy overall (%) by model, number of edges, and structure. Consistent with the feature importance [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Example model responses to a in-order traversal of a binary search tree. The extracted predictions are highlighted. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: A 3D perspective of the PCA image feature space from the first three components. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: The t-SNE feature space from the first two components. [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

75 extracted references · 58 canonical work pages

  1. [1]

    Lawrence Zitnick, Dhruv Batra, and Devi Parikh

    Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Dhruv Batra, and Devi Parikh. 2016. VQA: Visual Question Answering. arXiv:1505.00468 [cs.CL]

  2. [2]

    Shraddha Barke, Michael B James, and Nadia Polikarpova. 2022. Grounded Copilot: How Programmers Interact with Code-Generating Models.arXiv preprint arXiv:2206.15000 (2022)

  3. [3]

    Brett A Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos. 2023. Programming is hard-or at least it used to be: Educational opportunities and challenges of ai code generation. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education

  4. [4]

    Like a Nesting Doll

    Seth Bernstein, Paul Denny, Juho Leinonen, Lauren Kan, Arto Hellas, Matt Little- field, Sami Sarsa, and Stephen MacNeil. 2024. "Like a Nesting Doll": Analyzing Recursion Analogies Generated by CS Students Using Large Language Models. In Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1. 122–128

  5. [5]

    Seth Bernstein, Paul Denny, Juho Leinonen, Matt Littlefield, Arto Hellas, and Stephen MacNeil. 2024. Analyzing Students’ Preferences for LLM-Generated Analogies. In Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 2. 812–812

  6. [6]

    Susan Debra Blum. 2020. Ungrading: Why rating students undermines learning (and what to do instead) . West Virginia University Press

  7. [7]

    It’s not like Jarvis, but it’s pretty close!

    Ritvik Budhiraja, Ishika Joshi, Jagat Sesh Challa, Harshal D Akolekar, and Dhruv Kumar. 2024. “It’s not like Jarvis, but it’s pretty close!”-Examining ChatGPT’s Usage among Undergraduate Students in Computer Science. In Proceedings of the 26th Australasian Computing Education Conference . 124–133

  8. [8]

    Edy Budiman, Haeruddin Haeruddin, Ummul Hairah, and Faza Alameka. 2018. Mobile Learning: Visualizing Contents Media of Data Structures Course in Mobile Networks. Journal of Telecommunication, Electronic and Computer Engineering (JTEC) 10, 1-9 (Feb. 2018), 81–86. https://jtec.utem.edu.my/jtec/article/view/3877

Show all 75 references
  1. [9]

    Christopher Bull and Ahmed Kharrufa. 2023. Generative AI Assistants in Software Development Education: A vision for integrating Generative AI into educational practice, not instinctively defending against it

  2. [10]

    Michael D Byrne, Richard Catrambone, and John T Stasko. 1999. Evaluating ani- mations as student aids in learning computer algorithms. Computers & education 33, 4 (1999), 253–278

  3. [11]

    Michael B. Cahapay. 2021. Problems Encountered by College Students in Online Assessment Amid COVID-19 Crisis: A Case Study.SSRN Electronic Journal (2021)

  4. [12]

    Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps- Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, et al. 2023. MultiPL-E: a scalable and polyglot approach to benchmarking neural code generation. IEEE Transactio...

  5. [13]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374 (2021)

  6. [14]

    Bruno Pereira Cipriano and Pedro Alves. 2023. GPT-3 vs Object Oriented Program- ming Assignments: An Experience Report. In Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2023) . Association for Computing Machinery. ht...

  7. [15]

    Bruno Pereira Cipriano, Pedro Alves, and Paul Denny. 2024. A Picture Is Worth a Thousand Words: Exploring Diagram and Video-Based OOP Exercises to Counter LLM Over-Reliance. arXiv:2403.08396 [cs.SE]

  8. [16]

    Big Brother

    Simon Coghlan, Tim Miller, and Jeannie Paterson. 2021. Good Proctor or “Big Brother”? Ethics of Online Exam Supervision Technologies. Philosophy & Tech- nology 34, 4 (Aug. 2021), 1581–1606. https://doi.org/10.1007/s13347-021-00476-1

  9. [17]

    Rianne Conijn, Ad Kleingeld, Uwe Matzat, and Chris Snijders. 2022. The fear of big brother: The potential negative side-effects of proctored exams. Journal of Computer Assisted Learning 38, 6 (2022), 1521–1534

  10. [18]

    Vassilios Dagdilelis and Maria Satratzemi. 1998. DIDAGRAPH: Software for teaching graph theory algorithms. ACM SIGCSE Bulletin 30, 3 (1998), 64–68

  11. [19]

    Holger Danielsiek, Wolfgang Paul, and Jan Vahrenhold. 2012. Detecting and understanding students’ misconceptions related to algorithms and data structures. SIGCSE’12 - Proceedings of the 43rd ACM Technical Symposium on Computer Science Education (02 2012). https://doi.org/10.1...

  12. [20]

    Dwight Davis, J Kevin Dorsey, Ronald D Franks, Paul R Sackett, Cynthia A Searcy, and Xiaohui Zhao. 2013. Do racial and ethnic group differences in performance on the MCAT exam reflect test bias? Academic Medicine 88, 5 (2013), 593–602

  13. [21]

    Paul Denny, Viraj Kumar, and Nasser Giacaman. 2023. Conversing with Copilot: Exploring Prompt Engineering for Solving CS1 Problems Using Natural Lan- guage. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2023)

  14. [22]

    Becker, and Brent N

    Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, and Brent N. Reeves. 2023. Promptly: Using Prompt Problems to Teach Learners How to Effectively Utilize AI Code Generators. arXiv:2307.16364 [cs.HC]

  15. [23]

    Paul Denny, Stephen MacNeil, Jaromir Savelka, Leo Porter, and Andrew Luxton- Reilly. 2024. Desirable Characteristics for AI Teaching Assistants in Programming Education. In Proceedings of the 2024 on Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2024) . ...

  16. [24]

    Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan. 2023. Automated repair of programs from large language models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE

  17. [25]

    James Finnie-Ansley, Paul Denny, Andrew Luxton-Reilly, Eddie Antonio Santos, James Prather, and Brett A. Becker. 2023. My AI Wants to Know if This Will Be on the Exam: Testing OpenAI’s Codex on CS2 Programming Exercises. In Proceedings of the 25th Australasian Computing Educat...

  18. [26]

    Paul Gibson

    J. Paul Gibson. 2012. Teaching graph algorithms to children of all ages. In Proceedings of the 17th ACM Annual Conference on Innovation and Technology in Computer Science Education (ITiCSE ’12) . Association for Computing Machinery. ACE 2025, Gutierrez, et al

  19. [27]

    Golesteanu, Garrett B

    Matei A. Golesteanu, Garrett B. Vowinkel, and Ryan E. Dougherty. 2024. Can ChatGPT Pass a Theory of Computing Course? https://arxiv.org/abs/2407.07757

  20. [28]

    Carina S González-González, Alfonso Infante-Moro, and Juan C Infante-Moro

  21. [29]

    Scott Grissom, Laurie Murphy, Renée McCauley, and Sue Fitzgerald. 2016. Paper vs. Computer-based Exams: A Study of Errors in Recursive Binary Tree Algo- rithms. In Proceedings of the 47th ACM Technical Symposium on Computing Science Education (SIGCSE ’16). Association for Comp...

  22. [30]

    Sandra Gudiño Paredes, Felipe de Jesús Jasso Peña, and Juana María de La Fuente Alcazar. 2021. Remote proctored exams: Integrity assurance in online education? Distance Education 42, 2 (2021), 200–218

  23. [31]

    Neil Hatfield, Nathanial Brown, and Chad M Topaz. 2022. Do introductory courses disproportionately drive minoritized students out of STEM pathways? PNAS nexus 1, 4 (2022), pgac167

  24. [32]

    Jeffrey D Holmes. 2021. The Bad Test-Taker Identity. Teaching of psychology 48, 4 (2021), 293–299

  25. [34]

    Irene Hou, Sophia Mettille, Owen Man, Zhuo Li, Cynthia Zastudil, and Stephen MacNeil. 2024. The Effects of Generative AI on Introductory Students’ Help- Seeking Preferences. In Australasian Computing Education Conference (ACE ’24) . ACM. https://doi.org/10.1145/3636243.3636248

  26. [35]

    Association for Computing Machinery (ACM) Joint Task Force on Comput- ing Curricula and IEEE Computer Society. 2013. Computer Science Curricula 2013: Curriculum Guidelines for Undergraduate Degree Programs in Computer Science . Association for Computing Machinery, New York, NY, USA

  27. [36]

    Ishika Joshi, Ritvik Budhiraja, Harshal Dev, Jahnvi Kadia, Muhammad Osama, Ataullah, Sayan Mitra, and Dhruv Kumar. 2023. ChatGPT and the Future of Under- graduate Computer Science: Challenges, Opportunities and Recommendations. https://api.semanticscholar.org/CorpusID:258741417

  28. [37]

    Kuba Karpierz and Steven A. Wolfman. 2014. Misconceptions and concept inven- tory questions for binary search trees and hash tables. In Proceedings of the 45th ACM Technical Symposium on Computer Science Education (SIGCSE ’14) . Associa- tion for Computing Machinery, 6 pages. ...

  29. [38]

    Ericson, David Weintrop, and Tovi Grossman

    Majeed Kazemitabaar, Justin Chow, Carl Ka To Ma, Barbara J. Ericson, David Weintrop, and Tovi Grossman. 2023. Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming. In Proceedings of the 2023 CHI Conference on Human Factors in Comp...

  30. [39]

    Charles Koutcheme, Sami Sarsa, Juho Leinonen, Arto Hellas, and Paul Denny

  31. [40]

    Resistance is Futile

    Sam Lau and Philip J. Guo. 2023. From ‘Ban It Till We Understand It’ to "Resistance is Futile": How University Programming Instructors Plan to Adapt as More Students Use AI Code Generation and Explanation Tools such as ChatGPT and GitHub Copilot. In Proceedings of the 2023 ACM...

  32. [41]

    Juho Leinonen, Paul Denny, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, and Arto Hellas. 2023. Comparing Code Explanations Created by Students and Large Language Models. arXiv preprint arXiv:2304.03938 (2023)

  33. [42]

    Mark Liffiton, Brad Sheese, Jaromir Savelka, and Paul Denny. 2023. CodeHelp: Us- ing Large Language Models with Guardrails for Scalable Support in Programming Classes. arXiv:2308.06921 [cs.CY]

  34. [43]

    Stephen MacNeil, Paul Denny, Andrew Tran, Juho Leinonen, Seth Bernstein, Arto Hellas, Sami Sarsa, and Joanne Kim. 2023. Decoding Logic Errors: A Comparative Study on Bug Detection by Students and Large Language Models. arXiv preprint arXiv:2311.16017 (2023)

  35. [44]

    Becker, Michel Wermelinger, Arto Hellas, Andrew Tran, Sami Sarsa, James Prather, and Viraj Kumar

    Stephen MacNeil, Joanne Kim, Juho Leinonen, Paul Denny, Seth Bernstein, Brett A. Becker, Michel Wermelinger, Arto Hellas, Andrew Tran, Sami Sarsa, James Prather, and Viraj Kumar. 2023. The Implications of Large Language Models for CS Teachers and Students. In Proceedings of th...

  36. [45]

    Stephen MacNeil, Scott Spurlock, and Ian Applebaum. 2024. Imagining Comput- ing Education Assessment after Generative AI. https://arxiv.org/abs/2401.04601

  37. [46]

    Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny, Seth Bernstein, and Juho Leinonen. 2023. Experiences from Using Code Expla- nations Generated by Large Language Models in a Web Software Development E-Book. In Proc. SIGCSE’23. ACM, 6 pages

  38. [47]

    Stephen MacNeil, Andrew Tran, Dan Mogil, Seth Bernstein, Erin Ross, and Ziheng Huang. 2022. Generating Diverse Code Explanations Using the GPT-3 Large Language Model. In Proc. of the 2022 ACM Conf. on Int. Computing Education Research - Volume 2. ACM, 37–39

  39. [48]

    Janka Medová, Kitti Páleníková, Ľubomír Rybanský, and Zuzana Naštická. 2019. Undergraduate Students’ Solutions of Modeling Problems in Algorithmic Graph Theory. Mathematics 7, 7 (2019). https://doi.org/10.3390/math7070572

  40. [49]

    Laurie Murphy, Sue Fitzgerald, Scott Grissom, and Renée McCauley. 2015. Bug Infestation! A Goal-Plan Analysis of CS2 Students’ Recursive Binary Tree Solu- tions. In Proceedings of the 46th ACM Technical Symposium on Computer Science Education (SIGCSE ’15). Association for Comp...

  41. [50]

    Eng Lieh Ouh, Benjamin Kok Siew Gan, Kyong Jin Shim, and Swavek Wlodkowski

  42. [51]

    Ken Perlin, Zhenyi He, and Karl Rosenberg. 2018. Chalktalk : A Visualization and Communication Language – As a Tool in the Domain of Computer Science Education. arXiv:1809.07166 [cs.HC]

  43. [52]

    James Prather, Paul Denny, Juho Leinonen, Brett A Becker, Ibrahim Albluwi, Michael E Caspersen, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, et al . 2023. Transformed by Transformers: Navigating the AI Coding Revolution for Computing Education: An ITiCSE Workin...

  44. [53]

    In Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V

    ChatGPT, Can You Generate Solutions for my Coding Exercises? An Evaluation on its Effectiveness in an undergraduate Java Programming Course.. In Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V. 1. 54–60

  45. [54]

    James Prather, Brent N Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S Ran- drianasolo, Brett A Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Pro- grammers. In Proceedings of the 2024 ACM Confer...

  46. [55]

    Ben Puryear and Gina Sprint. 2022. Github copilot in the classroom: learning to code with AI assistance. Journal of Computing Sciences in Colleges 38, 1 (2022)

  47. [56]

    James Prather, Paul Denny, Juho Leinonen, Brett A Becker, Ibrahim Albluwi, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, Andrew Luxton- Reilly, et al. 2023. The Robots are Here: Navigating the Generative AI Revolution in Computing Education. arXiv preprint arXiv...

  48. [57]

    Ferozuddin Riaz and Khidir M. Ali. 2011. Applications of Graph Theory in Com- puter Science. In2011 Third International Conference on Computational Intelligence, Communication Systems and Networks . 142–145

  49. [58]

    Jürgen Rudolph, Samson Tan, and Shannon Tan. 2023. ChatGPT: Bullshit spewer or the end of traditional assessments in higher education? Journal of Applied Learning and Teaching 6, 1 (2023)

  50. [59]

    Brent Reeves, Sami Sarsa, James Prather, Paul Denny, Brett A Becker, Arto Hellas, Bailey Kimmel, Garrett Powell, and Juho Leinonen. 2023. Evaluating the perfor- mance of code generation models for solving Parsons problems with small prompt variations. In Proceedings of the 202...

  51. [60]

    Jaromir Savelka, Arav Agarwal, Christopher Bogart, and Majd Sakr. 2023. Large language models (gpt) struggle to answer multiple-choice questions about code. arXiv preprint arXiv:2303.08033 (2023)

  52. [61]

    Jaromir Savelka, Arav Agarwal, Christopher Bogart, Yifan Song, and Majd Sakr

  53. [62]

    Jaromir Savelka, Arav Agarwal, Marshall An, Chris Bogart, and Majd Sakr. 2023. Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses. The 19th ACM Conference on International Computing Education ...

  54. [63]

    Seidametova

    Zarema S. Seidametova. 2021. Some ways of increasing the efficiency of teaching data structures. In CoSinE@ICTERI

  55. [64]

    Teo Susnjak. 2022. ChatGPT: The End of Online Exam Integrity? arXiv:2212.09292

  56. [65]

    Can Generative Pre-trained Transformers (GPT) Pass Assessments in Higher Education Programming Courses? arXiv:2303.09325

  57. [66]

    Robert Sedgewick and Kevin Wayne. 2011. Algorithms. Pearson Education

  58. [67]

    Maarten Van Steen. 2010. Graph theory and complex networks. An introduction 144 (2010), 1–287

  59. [68]

    Michel Wermelinger. 2023. Using GitHub Copilot to Solve Simple Programming Problems. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2023). Association for Computing Machinery, New York, NY, USA, 172–178. https://doi.org/10.1145/35...

  60. [69]

    The Joint Task Force on Computing Curricula, Association for Computing Ma- chinery (ACM), IEEE Computer Society. 2013. Computer Science Curricula 2013: Curriculum Guidelines for Undergraduate Degree Programs in Computer Science. https://www.acm.org/binaries/content/assets/educ...

  61. [70]

    Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby, Jakob Verbeek, and Hervé Jégou. 2022. Three things everyone should know about vision transformers. In European Conference on Computer Vision . Springer

  62. [71]

    Daniel Zingaro, Cynthia Taylor, Leo Porter, Michael Clancy, Cynthia Lee, Soohyun Liao, and Kevin Webb. 2018. Identifying Student Difficulties with Basic Data Structures. 169–177. https://doi.org/10.1145/3230977.3231005

  63. [72]

    Andrea Ševčíková and Eva Milková. 2016. Multimedia applications: Graph algo- rithms visualization. In 2016 IEEE 17th International Symposium on Computational Intelligence and Informatics (CINTI) . 000231–000236. A EV ALUATION RESULTS In Table 2, we share additional analysis of...

  64. [73]

    Benjamin Xie, Matt J Davidson, Baker Franke, Emily McLeod, Min Li, and Amy J Ko. 2021. Domain Experts’ Interpretations of Assessment Bias in a Scaled, Online Computer Science Curriculum. In Proceedings of the Eighth ACM Conference on Learning@ Scale. 77–89

  65. [74]

    Cynthia Zastudil, Magdalena Rogalska, Christine Kapp, Jennifer Vaughn, and Stephen MacNeil. 2023. Generative AI in Computing Education: Perspectives of Students and Instructors. arXiv preprint arXiv:2308.04309 (2023). Seeing the Forest and the Trees ACE 2025,

  66. [2020]

    Sustainability 12, 8 (2020), 3488

    Implementation of e-proctoring in online teaching: A study about motiva- tional factors. Sustainability 12, 8 (2020), 3488

  67. [2023]

    In International Conference on Artificial Intelligence in Education

    Automated Program Repair Using Generative Models for Code Infilling. In International Conference on Artificial Intelligence in Education . Springer, 798–803

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.