Pith. sign in

REVIEW 2 major objections 1 minor 10 references

ExAtlas composes effects from nearby studies to link, reconcile or bridge social experiments

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-07-01 15:58 UTC pith:NFI5G554

load-bearing objection ExAtlas claims 98.6% recovery via local effect composition but only on targets pre-filtered by the same closeness metric used for the composition itself. the 2 major comments →

arxiv 2605.27153 v1 pith:NFI5G554 submitted 2026-05-26 cs.CY

Building an Atlas of Social Experiments to Link Studies, Reconcile Conflicts, and Bridge Gaps

classification cs.CY
keywords social experimentseffect compositionconflict reconciliationbridge experimentstreatment effect surfacelocal smoothnessatlas mapping
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

ExAtlas provides a way to build a structured atlas from the thousands of social experiments conducted each year. It locates studies that are close in treatment and outcome space and checks if their effects can be composed to match a new target study. Success with agreement creates links between studies, disagreement triggers reconciliation with possible moderators, and failure leads to suggestions for bridge experiments. This matters because it turns disconnected findings into a map that reveals consistent evidence, explains conflicts, and identifies what is missing, allowing knowledge to accumulate more effectively.

Core claim

The central discovery is that social experiment effects can be composed from locally close prior studies under a local smoothness assumption on the treatment-effect surface. Given a target, the method finds comparable studies, attempts composition, and classifies the outcome as a link if consistent, a conflict to reconcile if inconsistent, or a gap requiring a bridge experiment if composition is not possible. The approach includes an error bound for the composition. It achieves 98.6% recovery of effect direction on held-out locally supported targets, with human raters finding the bridge proposals plausible and the conflict explanations useful for generating theory.

What carries the argument

The composition operation on effects from studies close in treatment-outcome space, used to decide linking, reconciling, or bridging.

Load-bearing premise

Effects change smoothly enough in the local region of treatment and outcome space for composition from nearby studies to work with bounded error.

What would settle it

A test set of target studies where the composition either fails to recover the observed effect direction at rates much higher than 1.4 percent or where the proposed bridge experiments are judged implausible by domain experts.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Studies with consistent composed effects become linked in the atlas.
  • Inconsistent compositions lead to proposed moderators or theories to resolve conflicts.
  • Failed compositions identify gaps and suggest specific bridge experiments.
  • The experimental archive contains more latent structure than isolated studies reveal.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same composition approach might apply to experimental archives in other disciplines such as psychology or medicine.
  • Conflict explanations could feed into systems that propose and test new theories automatically.
  • Bridge suggestions could be used to design sequences of experiments that efficiently fill knowledge gaps.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper introduces ExAtlas, a framework for mapping social experiments into an atlas by retrieving locally close prior studies in treatment-outcome space. For a target study, it attempts to compose effects from neighbors: success with agreement yields a link to consistent evidence; success with disagreement yields conflict reconciliation via proposed moderators or theories; failure yields proposed bridge experiments. An error bound is derived under local smoothness of the treatment-effect surface. On held-out targets certified as locally supported, the method recovers effect direction in 98.6% of cases, and human evaluations indicate that proposed bridges are plausible and connected while conflict explanations aid theory generation.

Significance. If the evaluation concerns can be addressed, the framework could help accumulate and structure findings across the large volume of social and behavioral experiments, moving beyond isolated studies toward a more coherent map that guides both theory and new experiments. The explicit error bound under smoothness and the reported recovery rate represent concrete strengths in an area that often lacks such formalization.

major comments (2)
  1. [Abstract] Abstract: the 98.6% recovery rate is reported exclusively on held-out targets 'certified as locally supported.' Certification relies on the same local closeness metric used to retrieve and compose effects, so the result evaluates performance only on the subset where the local smoothness assumption is most likely to hold by selection; this does not test how often the assumption fails or how performance degrades outside the certified set.
  2. [Abstract] Abstract: the error bound for composition is derived under local smoothness, yet no results are provided on the fraction of targets that meet the local-support certification or on performance for non-certified targets; without these, the practical scope of the claimed recovery rate cannot be assessed.
minor comments (1)
  1. [Abstract] The abstract references human evaluations of bridge experiments and conflict explanations but provides no details on evaluator count, protocol, or agreement metrics; adding these would improve clarity of the supporting evidence.

Simulated Author's Rebuttal

2 responses · 0 unresolved

Thank you for the constructive review and for recognizing the potential of the framework along with its formal error bound and recovery results. We address the two major comments on the abstract evaluation below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the 98.6% recovery rate is reported exclusively on held-out targets 'certified as locally supported.' Certification relies on the same local closeness metric used to retrieve and compose effects, so the result evaluates performance only on the subset where the local smoothness assumption is most likely to hold by selection; this does not test how often the assumption fails or how performance degrades outside the certified set.

    Authors: The certification step is intentional: it isolates the regime in which the local smoothness assumption (and thus the derived error bound) is empirically supported by the presence of sufficiently close neighbors. Evaluating recovery only on this subset directly tests whether composition succeeds when the modeling assumptions hold, rather than averaging over cases where they do not. This is analogous to reporting in-distribution performance for a method whose guarantees are conditional. We agree that the fraction of targets meeting certification is needed to gauge how often the assumption is satisfied in practice, and we will add this statistic (computed on the held-out set) to the revised abstract and results section. Performance outside the certified set is not claimed to recover effects; non-certification instead triggers the bridge-experiment proposal, which is the appropriate output when local support is absent. revision: yes

  2. Referee: [Abstract] Abstract: the error bound for composition is derived under local smoothness, yet no results are provided on the fraction of targets that meet the local-support certification or on performance for non-certified targets; without these, the practical scope of the claimed recovery rate cannot be assessed.

    Authors: We will incorporate the fraction of held-out targets that satisfy the local-support certification criterion into the revised manuscript; this directly addresses the scope question. For non-certified targets the composition step is not executed, so a recovery rate is not applicable; the framework instead outputs a bridge-experiment proposal. The 98.6% figure is therefore presented with the explicit qualifier 'certified as locally supported' to avoid overgeneralization. If additional diagnostics on non-certified cases (e.g., frequency of certification failure) would strengthen the paper, we are prepared to include them. revision: yes

Circularity Check

0 steps flagged

No significant circularity detected

full rationale

The paper's core method composes effects from locally close studies under an explicit smoothness assumption for which an error bound is derived. The 98.6% direction-recovery figure is an empirical result on held-out targets, conditioned on a certification of local support that defines the regime where the composition is applicable. No equations or definitions in the provided text reduce the reported performance to the inputs by construction, no parameters are fitted and then relabeled as predictions, and no load-bearing self-citations or uniqueness theorems are invoked. The evaluation uses independent held-out studies and therefore remains externally falsifiable. This is a standard conditional-performance claim rather than a circular derivation.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The central claims rest on the domain assumption of local smoothness for error bounds and on the existence of a searchable archive of experiments with comparable treatment-outcome representations.

axioms (1)
  • domain assumption The treatment-effect surface is locally smooth
    Invoked to provide an error bound for composing effects from nearby studies.

pith-pipeline@v0.9.1-grok · 5784 in / 1180 out tokens · 30228 ms · 2026-07-01T15:58:22.197102+00:00 · methodology

0 comments
read the original abstract

Social and behavioral science runs thousands of experiments each year, yet their findings rarely accumulate into a coherent map of what is known, what conflicts, and what remains missing. We introduce ExAtlas, a framework for turning an archive of experiments into an atlas: a structured map in which studies link, conflict, or leave bridgeable gaps. Given a target study, ExAtlas searches for prior studies that are locally close in treatment and outcome space and asks whether their observed effects can be composed to predict the target effect. This yields three cases. If the composition succeeds and agrees with the observed result, ExAtlas links the target to consistent prior evidence. If composition succeeds but disagrees, ExAtlas reconciles the conflict and proposes candidate moderators or higher-level theories that could explain it. If composition fails, ExAtlas proposes bridge experiments to close the gap. We provide an error bound for composition under local smoothness of the treatment-effect surface. On held-out targets certified as locally supported, ExAtlas recovers effect direction in 98.6% of cases. Human evaluations further suggest that its proposed bridge experiments are plausible and exhibit connectedness, and that its conflict explanations are useful for theory generation. These results suggest that the archive of social experiments contains more latent structure than current practice extracts -- and that making this structure explicit can guide both future theory and future experimentation.

Figures

Figures reproduced from arXiv: 2605.27153 by Alex Yan, Honglin Bao, James A. Evans, Jiawei Zhang, Pengda Wang, Xiao Liu.

Figure 1
Figure 1. Figure 1: Illustration of how EXATLAS links, reconciles, and bridges experimental evidence. outcome pair. But the archive may contain studies of shorter mindfulness interventions before tasks, of study-skills supports targeted at first-generation students, and of brief contemplative practices in academic settings. If the target sits inside the re￾gion these studies span—close in treatment and outcome meanings—its ef… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the EXATLAS framework. Given a target experiment and an archive of prior experiments, EXATLAS tries to compose the target’s treatment and outcome from prior experiments. If the target is composable, EXATLAS estimates its effect from prior observed effects. Consistent estimates link experiments, mismatches motivate theoretical reconciliations, and non-composable targets trigger connecting experi… view at source ↗
Figure 3
Figure 3. Figure 3: DeepSeek-R1 dominates the arena when it comes to suggesting new experiments to connect exist￾ing ones (vs. OpenAI o3 and Claude Opus 4.5). C Evaluation Metrics for the Composition Experiment We evaluate predicted treatment effects using four metrics: sign match, MSE, MAE, and Spearman’s ρ. Let τˆt denote the predicted treatment effect for target experiment t, and let τt denote the reported human treatment … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

10 extracted references · 10 canonical work pages · 1 internal anchor

  1. [1]

    In Proceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Machine Learning Research, pages 337–371

    Using large language models to simulate mul- tiple humans and replicate human subject studies. In Proceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Machine Learning Research, pages 337–371. PMLR. Abdullah Almaatouq, Thomas L Griffiths, Jordan W Su- chow, Mark E Whiting, James Evans, and Duncan J Watts. 2024....

  2. [2]

    From Script to Stage: Automating Experimental Design for Social Simulations with LLMs

    Can ai language models replace human partici- pants?Trends in Cognitive Sciences, 27(7):597–600. Yuwei Guo, Zihan Zhao, Xiaowei Liu, Xiangning Yu, and Deyu Zhou. 2026. From script to stage: Au- tomating experimental design for social simulations with llms.Preprint, arXiv:2512.08935. 9 Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Robb Willer. 2024. ...

  3. [3]

    How does {IV} impact {DV}?

    Heuristic-based ideation for guid- ing llms toward structured creativity. InICLR Blogposts 2026. Https://iclr- blogposts.github.io/2026/blog/2026/ideation- heuristics/. Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Ya- mada, Shengran Hu, Jakob Foerster, David Ha, and Jeff Clune. 2026. Towards end-to-end automation of ai research.Nature, 651(8107):914–919...

  4. [4]

    Try to state the proposed experiment in the **simple form of independent/dependent variables and their relations** (e.g., ** Variable A positively/negatively impacts variable B.**)

  5. [5]

    You should try to **avoid** introducing complex conditioning, mediation, or moderation relations between dependent and independent variables unless doing so is necessary to establish a meaningful link between experiments

  6. [6]

    <**Avoid** introducing theories/experiments that directly compete with or contradict> the existing evidence or the derived synergy of existing evidence

  7. [7]

    Ensure the proposed experiments are ** logically coherent and empirically testable **

  8. [8]

    Embed **concrete details** into each proposed experiment

  9. [9]

    **Be concise and creative**

  10. [10]

    {listofknown}

    You should select the number of proposed experiments that is **the most appropriate** for the connection to the space of the established experiments. You should **avoid** proposing experiments that duplicate any of those in the given list below: "{listofknown}" Return only the proposed experiments. If more than one experiment is proposed, ** separate them...