REVIEW 11 cited by
Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We explore the abstract reasoning abilities of text-only and multimodal versions of GPT-4, using the ConceptARC benchmark [10], which is designed to evaluate robust understanding and reasoning with core-knowledge concepts. We extend the work of Moskvichev et al. [10] by evaluating GPT-4 on more detailed, one-shot prompting (rather than simple, zero-shot prompts) with text versions of ConceptARC tasks, and by evaluating GPT-4V, the multimodal version of GPT-4, on zero- and one-shot prompts using image versions of the simplest tasks. Our experimental results support the conclusion that neither version of GPT-4 has developed robust abstraction abilities at humanlike levels.
Forward citations
Cited by 11 Pith papers
-
Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction
Strict stage isolation that passes only a compressed symbolic schema and rule between LLM calls improves few-shot inductive reasoning more than self-refinement or explicit verbalization alone.
-
Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
Only the largest tested LLMs (about 70 billion parameters) match human accuracy on an abstract reasoning task, and the internal geometry of their best layers correlates moderately with human frontal EEG activity.
-
Dynamic Reinforcement Learning for Actors
A reinforcement learning update that adjusts each neuron's input-output sensitivity using TD error can replace external exploration noise and backpropagation through time in small actor-critic tasks.
-
Shuttle Between the Instructions and the Parameters of Large Language Models
SHIP jointly trains an encoder and decoder so a language model's soft-prompt parameters can be reconstructed from instructions and vice versa, improving instruction induction and inductive reasoning.
-
ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC)
ConceptSearch uses LLM-generated programs with concept-based scoring to solve 29/50 ARC training tasks and speed up search by up to 30% versus pixel-distance scoring.
-
What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models
Frozen LLMs reach near-perfect accuracy on several abstract reasoning benchmarks after tuning only the token embedding layer, implying poor zero-shot scores reflect input mismatch rather than absent reasoning.
-
Position: We Need An Algorithmic Understanding of Generative AI
The paper argues for a systematic algorithmic understanding of LLMs and presents a case study suggesting that Llama models do not implement BFS or DFS on graph navigation tasks.
-
NSA: Neuro-symbolic ARC Challenge
NSA, a neuro-symbolic ARC solver, solves 75 of 400 evaluation tasks by using a small transformer to propose DSL primitives that guide a combinatorial search.
-
The role of positional encodings in the ARC benchmark
In controlled experiments, 2D positional encoding outperforms 1D, RoPE, and learned embeddings for transformer models on ARC-like tasks when training examples are limited.
-
Impact of Noise on LLM-Models Performance in Abstraction and Reasoning Corpus (ARC) Tasks with Model Temperature Considerations
Adding tiny structured noise to ARC example grids sharply reduces GPT-4o's exact-match solve rate, but the paper's zero-noise success headline is built into its own task-selection rule.
-
MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree
MC-NEST adds a constant probability term to MCTSr's node selection and reports improved AIME pass@1 for GPT-4o, but the numbers are weakened by test-set rollout tuning and internal inconsistencies.
Discussion (0). Continue with ORCID to comment.