REVIEW 2 major objections 18 references
Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?
T0 review · 2 major / 0 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read Active exploration substantially improves adults' conjunctive causal reasoning in blicket tasks.
desk verdict Active exploration improves conjunctive performance over prior passive studies in this blicket setup, but the gain rests on cross-study comparison without a matched control, and LLM results stay preliminary. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Modified blicket detector task with unrestricted participant interventions to test causal structures.
What would settle it
A replication in which the same participants perform both active and passive versions of the identical interface and show no gain in conjunctive performance would falsify the improvement claim.
Extended reading notes
Core claim
Adult participants in an active blicket detector task identify conjunctive causal rules more effectively than in passive observation paradigms, closing much of the performance gap with disjunctive rules, although conjunctive inference still requires more interventions. State-of-the-art LLMs achieve comparable accuracy on rule inference but often use less efficient exploration and exhibit similar conjunctive-disjunctive gaps.
Load-bearing premise
The modified blicket detector task with unrestricted interventions measures pure causal inference ability without confounds from interface design, participant fatigue, or unmeasured strategy differences.
Editorial extensions
If this is right
- Active control over evidence generation helps overcome biases in causal learning.
- Conjunctive rules remain more demanding even when learners choose interventions.
- Some LLMs reach human-level accuracy on rule inference in this setting.
- Exploration strategies of LLMs differ from humans in efficiency.
- Task interface design affects measured causal reasoning performance.
Reading between the lines
- Simulating active exploration during LLM training could narrow efficiency gaps.
- The task setup could serve as a benchmark for causal abilities in other AI systems.
- Encouraging active testing in teaching might improve causal learning for students.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript examines whether granting adults agency via unrestricted interventions in a modified blicket detector task mitigates the well-documented conjunctive handicap in causal rule learning. It reports that active exploration yields substantially better conjunctive inference accuracy than prior passive-observation studies, although conjunctive rules still require more interventions to identify than disjunctive rules. The work also benchmarks a range of LLMs on the identical active task, finding that some frontier models reach near-human hypothesis accuracy while exhibiting less efficient exploration strategies and retaining similar conjunctive-disjunctive performance gaps.
Significance. If the active-exploration benefit is robustly demonstrated, the result would strengthen causal-learning accounts by showing that learner agency can overcome a persistent bias previously tied to passive evidence. The direct human-LLM comparison on an interventional task supplies a concrete benchmark for model scientific-reasoning capabilities and highlights differences in exploration efficiency that are not captured by accuracy alone.
major comments (2)
- [Abstract / Results] Abstract and §Results: the central claim that 'active exploration substantially improves adults' conjunctive causal reasoning' rests on a cross-study comparison to prior passive paradigms. No yoked passive-observation arm using the identical stimuli, interface, feedback timing, and rule structures is reported; without it, performance differences cannot be unambiguously attributed to agency rather than to unmeasured task-design factors.
- [Abstract] Abstract: no sample size, statistical tests, exclusion criteria, or precise task parameters (number of objects, intervention cost, stopping rule) are supplied, preventing evaluation of the data-to-claim mapping for the reported improvement and the conjunctive-disjunctive gap.
Simulated Author's Rebuttal
We thank the referee for their constructive comments on our manuscript. We respond to each major point below.
read point-by-point responses
-
Referee: [Abstract / Results] Abstract and §Results: the central claim that 'active exploration substantially improves adults' conjunctive causal reasoning' rests on a cross-study comparison to prior passive paradigms. No yoked passive-observation arm using the identical stimuli, interface, feedback timing, and rule structures is reported; without it, performance differences cannot be unambiguously attributed to agency rather than to unmeasured task-design factors.
Authors: We agree that a yoked passive-observation condition within the identical experimental setup would provide the most unambiguous evidence for the benefit of agency. Our design instead relies on comparison to established passive paradigms in the prior literature, with task parameters (objects, rules, feedback) aligned as closely as possible to those studies. While this leaves open the possibility that interface or timing differences contribute to the observed improvement, the magnitude of the gain relative to multiple prior reports supports the role of active exploration. We will add an explicit discussion of this limitation and the value of future matched controls in the revised manuscript. revision: partial
-
Referee: [Abstract] Abstract: no sample size, statistical tests, exclusion criteria, or precise task parameters (number of objects, intervention cost, stopping rule) are supplied, preventing evaluation of the data-to-claim mapping for the reported improvement and the conjunctive-disjunctive gap.
Authors: We will revise the abstract to report the sample size, statistical tests, exclusion criteria, number of objects, intervention cost, and stopping rule. These details will be added to strengthen the mapping from data to claims. revision: yes
Circularity Check
No circularity: purely empirical behavioral experiment with no derivations or fitted parameters
full rationale
The paper reports a new active-exploration blicket-detector experiment and compares results to prior passive-observation literature. No equations, parameter fitting, self-citation of uniqueness theorems, or ansatzes appear in the provided text. The central claim (active exploration improves conjunctive inference) is advanced via direct participant data rather than any reduction to the paper's own inputs or prior self-citations. Cross-study comparison is a methodological limitation but does not constitute circularity under the enumerated patterns.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?." pith.science (2026). https://pith.science/paper/SDE3DBVT
@misc{pith2026260606464,
author = {Pith},
title = {Pith review of: Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?},
year = {2026},
howpublished = {\url{https://pith.science/paper/SDE3DBVT}},
note = {Machine review of arXiv:2606.06464}
}
read the original abstract
A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the simultaneous presence of multiple causes, while performing better in disjunctive settings. However, most demonstrations of this ``conjunctive handicap'' rely on passive observation paradigms with limited evidence, where learners have no control over evidence generation. This paper asks whether this bias persists when adults are granted agency through active exploration. Using a modified ``blicket detector'' task, adult participants freely intervened to identify causal objects under conjunctive or disjunctive rule structures. We show that active exploration substantially improves adults' conjunctive causal reasoning, although conjunctive rules still require more tests to infer than disjunctive rules. We further compare human performance to a range of large language models in the same setting. While some state-of-the-art models approach human-level performance on hypothesis inference accuracy, they often exhibit less efficient exploration strategies and similar conjunctive-disjunctive performance gaps.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Child Development , volume=
Detecting blickets: How young children use information about novel causal powers in categorization and induction , author=. Child Development , volume=. 2000 , publisher=
2000
-
[2]
Psychological Science , volume=
When children are better (or at least more open-minded) learners than adults: Developmental differences in learning the forms of causal relationships , author=. Psychological Science , volume=. 2014 , publisher=
2014
-
[3]
Developmental Psychology , volume=
Preschool children learn about causal structure from conditional interventions , author=. Developmental Psychology , volume=. 2007 , publisher=
2007
-
[4]
2003 , publisher=
Making things happen: A theory of causal explanation , author=. 2003 , publisher=
2003
-
[5]
Psychological science , volume=
Design drives discovery in causal learning , author=. Psychological science , volume=. 2020 , publisher=
2020
-
[6]
Child development , volume=
Causal learning across culture and socioeconomic status , author=. Child development , volume=. 2019 , publisher=
2019
-
[7]
Cognition , volume=
Children are more exploratory and learn more than adults in an approach-avoid task , author=. Cognition , volume=. 2022 , publisher=
2022
-
[8]
Memory & Cognition , volume=
The importance of decision making in causal learning from interventions , author=. Memory & Cognition , volume=. 2006 , publisher=
2006
Show all 18 references
-
[9]
Proceedings of the 25th Annual Cognitive Science Society , pages=
Interventions do not solely benefit causal learning: Being told what to do results in worse learning than doing it yourself , author=. Proceedings of the 25th Annual Cognitive Science Society , pages=. 2013 , publisher=
2013
-
[10]
Advances in Neural Information Processing Systems , volume=
Doing experiments and revising rules with natural language and probabilistic reasoning , author=. Advances in Neural Information Processing Systems , volume=
-
[11]
Proceedings of the ieee/cvf conference on computer vision and pattern recognition , pages=
Acre: Abstract causal reasoning beyond covariation , author=. Proceedings of the ieee/cvf conference on computer vision and pattern recognition , pages=
-
[12]
How Can We Help Them Think Like Scientists? , author=
Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists? , author=. Second Conference on Language Modeling , year=
-
[13]
Proceedings of the National Academy of Sciences , volume=
Changes in cognitive flexibility and hypothesis search across human life history from childhood to adolescence to adulthood , author=. Proceedings of the National Academy of Sciences , volume=. 2017 , publisher=
2017
-
[14]
Cognition , volume=
Where science starts: Spontaneous experiments in preschoolers’ exploratory play , author=. Cognition , volume=. 2011 , publisher=
2011
-
[15]
Current directions in psychological science , volume=
When younger learners can be better (or at least more open-minded) than older ones , author=. Current directions in psychological science , volume=. 2015 , publisher=
2015
-
[16]
arXiv preprint arXiv:2512.08230 , year=
Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions , author=. arXiv preprint arXiv:2512.08230 , year=
-
[17]
Cognitive science , volume=
Children's causal inferences from indirect evidence: Backwards blocking and Bayesian reasoning in preschoolers , author=. Cognitive science , volume=. 2004 , publisher=
2004
-
[18]
Proceedings of the First Conference on Causal Learning and Reasoning , pages =
Learning Causal Overhypotheses through Exploration in Children and Computational Models , author =. Proceedings of the First Conference on Causal Learning and Reasoning , pages =. 2022 , editor =
2022
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.