On a newly generated benchmark, multimodal models solve up to 87.6% of visual tree problems and 56.2% of visual graph problems, undercutting the idea that diagrams make exam questions AI-proof.
Can ChatGPT Pass a Theory of Computing Course?
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Large Language Models (LLMs) have had considerable difficulty when prompted with mathematical questions, especially those within theory of computing (ToC) courses. In this paper, we detail two experiments regarding our own ToC course and the ChatGPT LLM. For the first, we evaluated ChatGPT's ability to pass our own ToC course's exams. For the second, we created a database of sample ToC questions and responses to accommodate other ToC offerings' choices for topics and structure. We scored each of ChatGPT's outputs on these questions. Overall, we determined that ChatGPT can pass our ToC course, and is adequate at understanding common formal definitions and answering "simple"-style questions, e.g., true/false and multiple choice. However, ChatGPT often makes nonsensical claims in open-ended responses, such as proofs.
citation-role summary
citation-polarity summary
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Seeing the Forest and the Trees: Solving Visual Graph and Tree Based Data Structure Problems using Large Multimodal Models
On a newly generated benchmark, multimodal models solve up to 87.6% of visual tree problems and 56.2% of visual graph problems, undercutting the idea that diagrams make exam questions AI-proof.