REVIEW 6 cited by
Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Human language offers a powerful window into our thoughts -- we tell stories, give explanations, and express our beliefs and goals through words. Abundant evidence also suggests that language plays a developmental role in structuring our learning. Here, we ask: how much of human-like thinking can be captured by learning statistical patterns in language alone? We first contribute a new challenge benchmark for comparing humans and distributional large language models (LLMs). Our benchmark contains two problem-solving domains (planning and explanation generation) and is designed to require generalization to new, out-of-distribution problems expressed in language. We find that humans are far more robust than LLMs on this benchmark. Next, we propose a hybrid Parse-and-Solve model, which augments distributional LLMs with a structured symbolic reasoning module. We find that this model shows more robust adaptation to out-of-distribution planning problems, demonstrating the promise of hybrid AI models for more human-like reasoning.
Forward citations
Cited by 6 Pith papers
-
Can Vision Language Models Learn Intuitive Physics from Interaction?
Training VLMs through interaction (GRPO) does not yield generalizable physical intuitions beyond within-task performance, matching—not exceeding—supervised fine-tuning.
-
Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
GUI agents frequently fall for deceptive interface designs, often without recognizing them, and human supervision of agents improves avoidance only partially while introducing new attention and workload costs.
-
Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
Wrapping LLM calls as oracles in classical algorithms improves sorting and clustering accuracy for small models, but several advertised applications are missing from the experiments.
-
Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering
A neuro-symbolic Rational Speech Act model with LLM proposers and evaluators predicts human question-answer patterns about as well as the fully hand-specified probabilistic model.
-
Annotating Compositionality Scores for Irish Noun Compounds is Hard Work
The authors present annotation guidelines and a 270-item pilot corpus of Irish noun compounds with compositionality and related scores, but the dataset itself is not yet released.
-
THiNK: Can Large Language Models Think-aloud?
THiNK uses a multi-agent, feedback-driven loop of problem revision and GPT-4O-based Bloom's Taxonomy scoring to measure and improve higher-order thinking in LLMs on math word problems.
Discussion (0). Sign in to comment.