Pith. sign in

REVIEW 2 major objections 18 references

Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?

T0 review · 2 major / 0 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read Active exploration substantially improves adults' conjunctive causal reasoning in blicket tasks.

desk verdict Active exploration improves conjunctive performance over prior passive studies in this blicket setup, but the gain rests on cross-study comparison without a matched control, and LLM results stay preliminary. read the letter →

arxiv 2606.06464 v1 pith:SDE3DBVT submitted 2026-06-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords causalreasoningactiveexplorationconjunctiverulesblicketdetectorlargelanguagemodelshumanperformancecomparison
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests whether adults' known difficulty identifying conjunctive causal rules persists when they can actively generate evidence rather than observe passively. In a modified blicket detector, participants freely intervene to discover which objects activate a machine under conjunctive or disjunctive rules. Results indicate active exploration raises conjunctive performance well above prior passive benchmarks, yet conjunctive rules still take more tests to infer than disjunctive ones. The same tasks applied to LLMs show some models reach near-human inference accuracy while using less efficient exploration strategies and retaining similar rule-type gaps.

What carries the argument

Modified blicket detector task with unrestricted participant interventions to test causal structures.

What would settle it

A replication in which the same participants perform both active and passive versions of the identical interface and show no gain in conjunctive performance would falsify the improvement claim.

Watch

Extended reading notes

Core claim

Adult participants in an active blicket detector task identify conjunctive causal rules more effectively than in passive observation paradigms, closing much of the performance gap with disjunctive rules, although conjunctive inference still requires more interventions. State-of-the-art LLMs achieve comparable accuracy on rule inference but often use less efficient exploration and exhibit similar conjunctive-disjunctive gaps.

Load-bearing premise

The modified blicket detector task with unrestricted interventions measures pure causal inference ability without confounds from interface design, participant fatigue, or unmeasured strategy differences.

Editorial extensions

If this is right

  • Active control over evidence generation helps overcome biases in causal learning.
  • Conjunctive rules remain more demanding even when learners choose interventions.
  • Some LLMs reach human-level accuracy on rule inference in this setting.
  • Exploration strategies of LLMs differ from humans in efficiency.
  • Task interface design affects measured causal reasoning performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Simulating active exploration during LLM training could narrow efficiency gaps.
  • The task setup could serve as a benchmark for causal abilities in other AI systems.
  • Encouraging active testing in teaching might improve causal learning for students.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The manuscript examines whether granting adults agency via unrestricted interventions in a modified blicket detector task mitigates the well-documented conjunctive handicap in causal rule learning. It reports that active exploration yields substantially better conjunctive inference accuracy than prior passive-observation studies, although conjunctive rules still require more interventions to identify than disjunctive rules. The work also benchmarks a range of LLMs on the identical active task, finding that some frontier models reach near-human hypothesis accuracy while exhibiting less efficient exploration strategies and retaining similar conjunctive-disjunctive performance gaps.

Significance. If the active-exploration benefit is robustly demonstrated, the result would strengthen causal-learning accounts by showing that learner agency can overcome a persistent bias previously tied to passive evidence. The direct human-LLM comparison on an interventional task supplies a concrete benchmark for model scientific-reasoning capabilities and highlights differences in exploration efficiency that are not captured by accuracy alone.

major comments (2)
  1. [Abstract / Results] Abstract and §Results: the central claim that 'active exploration substantially improves adults' conjunctive causal reasoning' rests on a cross-study comparison to prior passive paradigms. No yoked passive-observation arm using the identical stimuli, interface, feedback timing, and rule structures is reported; without it, performance differences cannot be unambiguously attributed to agency rather than to unmeasured task-design factors.
  2. [Abstract] Abstract: no sample size, statistical tests, exclusion criteria, or precise task parameters (number of objects, intervention cost, stopping rule) are supplied, preventing evaluation of the data-to-claim mapping for the reported improvement and the conjunctive-disjunctive gap.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive comments on our manuscript. We respond to each major point below.

read point-by-point responses
  1. Referee: [Abstract / Results] Abstract and §Results: the central claim that 'active exploration substantially improves adults' conjunctive causal reasoning' rests on a cross-study comparison to prior passive paradigms. No yoked passive-observation arm using the identical stimuli, interface, feedback timing, and rule structures is reported; without it, performance differences cannot be unambiguously attributed to agency rather than to unmeasured task-design factors.

    Authors: We agree that a yoked passive-observation condition within the identical experimental setup would provide the most unambiguous evidence for the benefit of agency. Our design instead relies on comparison to established passive paradigms in the prior literature, with task parameters (objects, rules, feedback) aligned as closely as possible to those studies. While this leaves open the possibility that interface or timing differences contribute to the observed improvement, the magnitude of the gain relative to multiple prior reports supports the role of active exploration. We will add an explicit discussion of this limitation and the value of future matched controls in the revised manuscript. revision: partial

  2. Referee: [Abstract] Abstract: no sample size, statistical tests, exclusion criteria, or precise task parameters (number of objects, intervention cost, stopping rule) are supplied, preventing evaluation of the data-to-claim mapping for the reported improvement and the conjunctive-disjunctive gap.

    Authors: We will revise the abstract to report the sample size, statistical tests, exclusion criteria, number of objects, intervention cost, and stopping rule. These details will be added to strengthen the mapping from data to claims. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: purely empirical behavioral experiment with no derivations or fitted parameters

full rationale

The paper reports a new active-exploration blicket-detector experiment and compares results to prior passive-observation literature. No equations, parameter fitting, self-citation of uniqueness theorems, or ansatzes appear in the provided text. The central claim (active exploration improves conjunctive inference) is advanced via direct participant data rather than any reduction to the paper's own inputs or prior self-citations. Cross-study comparison is a methodological limitation but does not constitute circularity under the enumerated patterns.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Empirical study; no mathematical free parameters, axioms, or invented entities are stated in the abstract. The design implicitly assumes the blicket detector interface validly isolates causal inference.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?." pith.science (2026). https://pith.science/paper/SDE3DBVT

@misc{pith2026260606464,
  author       = {Pith},
  title        = {Pith review of: Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SDE3DBVT}},
  note         = {Machine review of arXiv:2606.06464}
}
read the original abstract

A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the simultaneous presence of multiple causes, while performing better in disjunctive settings. However, most demonstrations of this ``conjunctive handicap'' rely on passive observation paradigms with limited evidence, where learners have no control over evidence generation. This paper asks whether this bias persists when adults are granted agency through active exploration. Using a modified ``blicket detector'' task, adult participants freely intervened to identify causal objects under conjunctive or disjunctive rule structures. We show that active exploration substantially improves adults' conjunctive causal reasoning, although conjunctive rules still require more tests to infer than disjunctive rules. We further compare human performance to a range of large language models in the same setting. While some state-of-the-art models approach human-level performance on hypothesis inference accuracy, they often exhibit less efficient exploration strategies and similar conjunctive-disjunctive performance gaps.

Figures

Figures reproduced from arXiv: 2606.06464 by the authors.

Figure 1
Figure 1. Test structure in the Active Exploration vs. Passive Observation conditions. In the active exploration condition, participants can click to add or remove four individual objects from the Nexiom machine, and explicitly test the current combination to observe whether the machine switches ON or OFF. In the passive observation condition, participants are each paired to an active exploration participant (yoked control). … view at source ↗
Figure 2
Figure 2. Rule inference accuracy across age groups. Performance of children and adults when given ambiguous passive-observation data in a prior study (Gopnik et al. (2017) (left of dashed line) is compared with adults’ performance in the current study (right of dashed line). to the Passive Observation condition, and another 102 were as￾signed to the Passive Proposer condition. Participants in both passive conditions were ran… view at source ↗
Figure 3
Figure 3. Performance accuracy in conjunctive vs. dis￾junctive causal inference with active exploration evidence. Error bars are ± standard error of the mean across participants. Results Beyond overall accuracy, our results show that both human participants and language models do not interact with the Nex￾iom machine in a random or unguided manner. Instead, they engaged in systematic exploration designed to disambiguate compe… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4 [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: Sequential search strategy across human and language agents during active exploration. Each color rep￾resents the dominant number of objects per test index, with opacity indicating the percentage of trials that have that dom￾inant number of objects within that index. G…
Figure 5
Figure 5. Figure 5: Distribution of rule and object identification outcomes for conjunctive and disjunctive conditions in adults and LLMs with active exploration. Complexity of Hypothesis Space and Search Strategy. The analysis of information gain patterns reveals systematic differences i…
Figure 7
Figure 7. Figure 7 [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

18 extracted references · 1 canonical work pages

  1. [1]

    Child Development , volume=

    Detecting blickets: How young children use information about novel causal powers in categorization and induction , author=. Child Development , volume=. 2000 , publisher=

  2. [2]

    Psychological Science , volume=

    When children are better (or at least more open-minded) learners than adults: Developmental differences in learning the forms of causal relationships , author=. Psychological Science , volume=. 2014 , publisher=

  3. [3]

    Developmental Psychology , volume=

    Preschool children learn about causal structure from conditional interventions , author=. Developmental Psychology , volume=. 2007 , publisher=

  4. [4]

    2003 , publisher=

    Making things happen: A theory of causal explanation , author=. 2003 , publisher=

  5. [5]

    Psychological science , volume=

    Design drives discovery in causal learning , author=. Psychological science , volume=. 2020 , publisher=

  6. [6]

    Child development , volume=

    Causal learning across culture and socioeconomic status , author=. Child development , volume=. 2019 , publisher=

  7. [7]

    Cognition , volume=

    Children are more exploratory and learn more than adults in an approach-avoid task , author=. Cognition , volume=. 2022 , publisher=

  8. [8]

    Memory & Cognition , volume=

    The importance of decision making in causal learning from interventions , author=. Memory & Cognition , volume=. 2006 , publisher=

Show all 18 references
  1. [9]

    Proceedings of the 25th Annual Cognitive Science Society , pages=

    Interventions do not solely benefit causal learning: Being told what to do results in worse learning than doing it yourself , author=. Proceedings of the 25th Annual Cognitive Science Society , pages=. 2013 , publisher=

  2. [10]

    Advances in Neural Information Processing Systems , volume=

    Doing experiments and revising rules with natural language and probabilistic reasoning , author=. Advances in Neural Information Processing Systems , volume=

  3. [11]

    Proceedings of the ieee/cvf conference on computer vision and pattern recognition , pages=

    Acre: Abstract causal reasoning beyond covariation , author=. Proceedings of the ieee/cvf conference on computer vision and pattern recognition , pages=

  4. [12]

    How Can We Help Them Think Like Scientists? , author=

    Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists? , author=. Second Conference on Language Modeling , year=

  5. [13]

    Proceedings of the National Academy of Sciences , volume=

    Changes in cognitive flexibility and hypothesis search across human life history from childhood to adolescence to adulthood , author=. Proceedings of the National Academy of Sciences , volume=. 2017 , publisher=

  6. [14]

    Cognition , volume=

    Where science starts: Spontaneous experiments in preschoolers’ exploratory play , author=. Cognition , volume=. 2011 , publisher=

  7. [15]

    Current directions in psychological science , volume=

    When younger learners can be better (or at least more open-minded) than older ones , author=. Current directions in psychological science , volume=. 2015 , publisher=

  8. [16]

    arXiv preprint arXiv:2512.08230 , year=

    Empowerment Gain and Causal Model Construction: Children and adults are sensitive to controllability and variability in their causal interventions , author=. arXiv preprint arXiv:2512.08230 , year=

  9. [17]

    Cognitive science , volume=

    Children's causal inferences from indirect evidence: Backwards blocking and Bayesian reasoning in preschoolers , author=. Cognitive science , volume=. 2004 , publisher=

  10. [18]

    Proceedings of the First Conference on Causal Learning and Reasoning , pages =

    Learning Causal Overhypotheses through Exploration in Children and Computational Models , author =. Proceedings of the First Conference on Causal Learning and Reasoning , pages =. 2022 , editor =

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.