REVIEW 2 major objections 1 minor 10 references
ExAtlas composes effects from nearby studies to link, reconcile or bridge social experiments
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-07-01 15:58 UTC pith:NFI5G554
load-bearing objection ExAtlas claims 98.6% recovery via local effect composition but only on targets pre-filtered by the same closeness metric used for the composition itself. the 2 major comments →
Building an Atlas of Social Experiments to Link Studies, Reconcile Conflicts, and Bridge Gaps
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that social experiment effects can be composed from locally close prior studies under a local smoothness assumption on the treatment-effect surface. Given a target, the method finds comparable studies, attempts composition, and classifies the outcome as a link if consistent, a conflict to reconcile if inconsistent, or a gap requiring a bridge experiment if composition is not possible. The approach includes an error bound for the composition. It achieves 98.6% recovery of effect direction on held-out locally supported targets, with human raters finding the bridge proposals plausible and the conflict explanations useful for generating theory.
What carries the argument
The composition operation on effects from studies close in treatment-outcome space, used to decide linking, reconciling, or bridging.
Load-bearing premise
Effects change smoothly enough in the local region of treatment and outcome space for composition from nearby studies to work with bounded error.
What would settle it
A test set of target studies where the composition either fails to recover the observed effect direction at rates much higher than 1.4 percent or where the proposed bridge experiments are judged implausible by domain experts.
If this is right
- Studies with consistent composed effects become linked in the atlas.
- Inconsistent compositions lead to proposed moderators or theories to resolve conflicts.
- Failed compositions identify gaps and suggest specific bridge experiments.
- The experimental archive contains more latent structure than isolated studies reveal.
Where Pith is reading between the lines
- The same composition approach might apply to experimental archives in other disciplines such as psychology or medicine.
- Conflict explanations could feed into systems that propose and test new theories automatically.
- Bridge suggestions could be used to design sequences of experiments that efficiently fill knowledge gaps.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ExAtlas, a framework for mapping social experiments into an atlas by retrieving locally close prior studies in treatment-outcome space. For a target study, it attempts to compose effects from neighbors: success with agreement yields a link to consistent evidence; success with disagreement yields conflict reconciliation via proposed moderators or theories; failure yields proposed bridge experiments. An error bound is derived under local smoothness of the treatment-effect surface. On held-out targets certified as locally supported, the method recovers effect direction in 98.6% of cases, and human evaluations indicate that proposed bridges are plausible and connected while conflict explanations aid theory generation.
Significance. If the evaluation concerns can be addressed, the framework could help accumulate and structure findings across the large volume of social and behavioral experiments, moving beyond isolated studies toward a more coherent map that guides both theory and new experiments. The explicit error bound under smoothness and the reported recovery rate represent concrete strengths in an area that often lacks such formalization.
major comments (2)
- [Abstract] Abstract: the 98.6% recovery rate is reported exclusively on held-out targets 'certified as locally supported.' Certification relies on the same local closeness metric used to retrieve and compose effects, so the result evaluates performance only on the subset where the local smoothness assumption is most likely to hold by selection; this does not test how often the assumption fails or how performance degrades outside the certified set.
- [Abstract] Abstract: the error bound for composition is derived under local smoothness, yet no results are provided on the fraction of targets that meet the local-support certification or on performance for non-certified targets; without these, the practical scope of the claimed recovery rate cannot be assessed.
minor comments (1)
- [Abstract] The abstract references human evaluations of bridge experiments and conflict explanations but provides no details on evaluator count, protocol, or agreement metrics; adding these would improve clarity of the supporting evidence.
Simulated Author's Rebuttal
Thank you for the constructive review and for recognizing the potential of the framework along with its formal error bound and recovery results. We address the two major comments on the abstract evaluation below.
read point-by-point responses
-
Referee: [Abstract] Abstract: the 98.6% recovery rate is reported exclusively on held-out targets 'certified as locally supported.' Certification relies on the same local closeness metric used to retrieve and compose effects, so the result evaluates performance only on the subset where the local smoothness assumption is most likely to hold by selection; this does not test how often the assumption fails or how performance degrades outside the certified set.
Authors: The certification step is intentional: it isolates the regime in which the local smoothness assumption (and thus the derived error bound) is empirically supported by the presence of sufficiently close neighbors. Evaluating recovery only on this subset directly tests whether composition succeeds when the modeling assumptions hold, rather than averaging over cases where they do not. This is analogous to reporting in-distribution performance for a method whose guarantees are conditional. We agree that the fraction of targets meeting certification is needed to gauge how often the assumption is satisfied in practice, and we will add this statistic (computed on the held-out set) to the revised abstract and results section. Performance outside the certified set is not claimed to recover effects; non-certification instead triggers the bridge-experiment proposal, which is the appropriate output when local support is absent. revision: yes
-
Referee: [Abstract] Abstract: the error bound for composition is derived under local smoothness, yet no results are provided on the fraction of targets that meet the local-support certification or on performance for non-certified targets; without these, the practical scope of the claimed recovery rate cannot be assessed.
Authors: We will incorporate the fraction of held-out targets that satisfy the local-support certification criterion into the revised manuscript; this directly addresses the scope question. For non-certified targets the composition step is not executed, so a recovery rate is not applicable; the framework instead outputs a bridge-experiment proposal. The 98.6% figure is therefore presented with the explicit qualifier 'certified as locally supported' to avoid overgeneralization. If additional diagnostics on non-certified cases (e.g., frequency of certification failure) would strengthen the paper, we are prepared to include them. revision: yes
Circularity Check
No significant circularity detected
full rationale
The paper's core method composes effects from locally close studies under an explicit smoothness assumption for which an error bound is derived. The 98.6% direction-recovery figure is an empirical result on held-out targets, conditioned on a certification of local support that defines the regime where the composition is applicable. No equations or definitions in the provided text reduce the reported performance to the inputs by construction, no parameters are fitted and then relabeled as predictions, and no load-bearing self-citations or uniqueness theorems are invoked. The evaluation uses independent held-out studies and therefore remains externally falsifiable. This is a standard conditional-performance claim rather than a circular derivation.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption The treatment-effect surface is locally smooth
read the original abstract
Social and behavioral science runs thousands of experiments each year, yet their findings rarely accumulate into a coherent map of what is known, what conflicts, and what remains missing. We introduce ExAtlas, a framework for turning an archive of experiments into an atlas: a structured map in which studies link, conflict, or leave bridgeable gaps. Given a target study, ExAtlas searches for prior studies that are locally close in treatment and outcome space and asks whether their observed effects can be composed to predict the target effect. This yields three cases. If the composition succeeds and agrees with the observed result, ExAtlas links the target to consistent prior evidence. If composition succeeds but disagrees, ExAtlas reconciles the conflict and proposes candidate moderators or higher-level theories that could explain it. If composition fails, ExAtlas proposes bridge experiments to close the gap. We provide an error bound for composition under local smoothness of the treatment-effect surface. On held-out targets certified as locally supported, ExAtlas recovers effect direction in 98.6% of cases. Human evaluations further suggest that its proposed bridge experiments are plausible and exhibit connectedness, and that its conflict explanations are useful for theory generation. These results suggest that the archive of social experiments contains more latent structure than current practice extracts -- and that making this structure explicit can guide both future theory and future experimentation.
Figures
Reference graph
Works this paper leans on
-
[1]
Using large language models to simulate mul- tiple humans and replicate human subject studies. In Proceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Machine Learning Research, pages 337–371. PMLR. Abdullah Almaatouq, Thomas L Griffiths, Jordan W Su- chow, Mark E Whiting, James Evans, and Duncan J Watts. 2024....
work page 2024
-
[2]
From Script to Stage: Automating Experimental Design for Social Simulations with LLMs
Can ai language models replace human partici- pants?Trends in Cognitive Sciences, 27(7):597–600. Yuwei Guo, Zihan Zhao, Xiaowei Liu, Xiangning Yu, and Deyu Zhou. 2026. From script to stage: Au- tomating experimental design for social simulations with llms.Preprint, arXiv:2512.08935. 9 Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Robb Willer. 2024. ...
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[3]
Heuristic-based ideation for guid- ing llms toward structured creativity. InICLR Blogposts 2026. Https://iclr- blogposts.github.io/2026/blog/2026/ideation- heuristics/. Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Ya- mada, Shengran Hu, Jakob Foerster, David Ha, and Jeff Clune. 2026. Towards end-to-end automation of ai research.Nature, 651(8107):914–919...
-
[4]
Try to state the proposed experiment in the **simple form of independent/dependent variables and their relations** (e.g., ** Variable A positively/negatively impacts variable B.**)
-
[5]
You should try to **avoid** introducing complex conditioning, mediation, or moderation relations between dependent and independent variables unless doing so is necessary to establish a meaningful link between experiments
-
[6]
<**Avoid** introducing theories/experiments that directly compete with or contradict> the existing evidence or the derived synergy of existing evidence
-
[7]
Ensure the proposed experiments are ** logically coherent and empirically testable **
-
[8]
Embed **concrete details** into each proposed experiment
-
[9]
**Be concise and creative**
-
[10]
You should select the number of proposed experiments that is **the most appropriate** for the connection to the space of the established experiments. You should **avoid** proposing experiments that duplicate any of those in the given list below: "{listofknown}" Return only the proposed experiments. If more than one experiment is proposed, ** separate them...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.