Pith. sign in

REVIEW 6 cited by

Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.05718 v1 pith:FJZ27UIC submitted 2022-05-11 cs.CL cs.AIcs.LGcs.SC

classification cs.CLcs.AIcs.LGcs.SC
keywords languagebenchmarkhuman-likellmsmodelsout-of-distributionreasoningrobust
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human language offers a powerful window into our thoughts -- we tell stories, give explanations, and express our beliefs and goals through words. Abundant evidence also suggests that language plays a developmental role in structuring our learning. Here, we ask: how much of human-like thinking can be captured by learning statistical patterns in language alone? We first contribute a new challenge benchmark for comparing humans and distributional large language models (LLMs). Our benchmark contains two problem-solving domains (planning and explanation generation) and is designed to require generalization to new, out-of-distribution problems expressed in language. We find that humans are far more robust than LLMs on this benchmark. Next, we propose a hybrid Parse-and-Solve model, which augments distributional LLMs with a structured symbolic reasoning module. We find that this model shows more robust adaptation to out-of-distribution planning problems, demonstrating the promise of hybrid AI models for more human-like reasoning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. Can Vision Language Models Learn Intuitive Physics from Interaction?

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Training VLMs through interaction (GRPO) does not yield generalizable physical intuitions beyond within-task performance, matching—not exceeding—supervised fine-tuning.

  2. Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight

    cs.HC 2025-09 conditional novelty 6.0 of 10

    GUI agents frequently fall for deceptive interface designs, often without recognizing them, and human supervision of agents improves avoidance only partially while introducing new attention and workload costs.

  3. Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)

    cs.CL 2025-09 unverdicted novelty 6.0 of 10

    Wrapping LLM calls as oracles in classical algorithms improves sorting and clustering accuracy for small models, but several advertised applications are missing from the experiments.

  4. Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A neuro-symbolic Rational Speech Act model with LLM proposers and evaluators predicts human question-answer patterns about as well as the fully hand-specified probabilistic model.

  5. Annotating Compositionality Scores for Irish Noun Compounds is Hard Work

    cs.CL 2025-02 conditional novelty 6.0 of 10

    The authors present annotation guidelines and a 270-item pilot corpus of Irish noun compounds with compositionality and related scores, but the dataset itself is not yet released.

  6. THiNK: Can Large Language Models Think-aloud?

    cs.CL 2025-05 reject novelty 4.0 of 10

    THiNK uses a multi-agent, feedback-driven loop of problem revision and GPT-4O-based Bloom's Taxonomy scoring to measure and improve higher-order thinking in LLMs on math word problems.

Pith tools