{"id":"eeed0738-0a8d-4b13-bd8e-9206d0d245df","arxiv_id":"2605.27153","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ExAtlas composes effects from locally close prior studies in treatment-outcome space to link, reconcile conflicts, or propose bridge experiments, recovering effect direction in 98.6% of held-out locally supported targets under local smoothness.","lead":"The paper introduces ExAtlas, a framework that maps social experiments by searching for nearby studies in treatment and outcome space and composing their effects to predict, reconcile with, or bridge gaps to a target study. A smart generalist might read it to understand how thousands of isolated behavioral experiments could be turned into a connected, navigable knowledge structure.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"98.6% recovery reported only on 'locally supported' held-out targets whose certification uses the same closeness metric as composition","rationale":"Reader correctly flags the smoothness assumption as central to the bound, but the load-bearing issue for the quoted 98.6% claim is the certification filter that preconditions the evaluation on the assumption holding. Full text would be required to confirm the selection procedure, but the abstract alone makes this the clearest internal risk to the performance claim. No other inconsistency is visible from supplied material.","tokens_in":1756,"tokens_out":330,"duration_ms":23299,"concrete_test":"From the methods section, extract the exact definition and distance threshold for 'locally supported' certification; recompute direction recovery on the full held-out set (or a random subsample) without the filter and with a 2× larger distance threshold; if accuracy drops below 85% or the certified subset is <30% of total held-out targets, the headline figure is not representative.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim conditions performance on targets 'certified as locally supported.' Certification appears to select cases with nearby studies in treatment-outcome space—the identical criterion used to retrieve and compose effects. This creates a selection filter that excludes precisely the regimes where local smoothness fails or no close neighbors exist. The error bound is derived under the smoothness assumption, but the reported 98.6% does not test how often that assumption holds or how performance degrades outside the certified subset. Human evaluations of bridges and conflicts are downstream of this filtered set.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces ExAtlas, a framework for mapping social experiments into an atlas by retrieving locally close prior studies in treatment-outcome space. For a target study, it attempts to compose effects from neighbors: success with agreement yields a link to consistent evidence; success with disagreement yields conflict reconciliation via proposed moderators or theories; failure yields proposed bridge experiments. An error bound is derived under local smoothness of the treatment-effect surface. On held-out targets certified as locally supported, the method recovers effect direction in 98.6% of cases, and human evaluations indicate that proposed bridges are plausible and connected while conflict explanations aid theory generation.","tokens_in":1882,"tokens_out":484,"duration_ms":24344,"significance":"If the evaluation concerns can be addressed, the framework could help accumulate and structure findings across the large volume of social and behavioral experiments, moving beyond isolated studies toward a more coherent map that guides both theory and new experiments. The explicit error bound under smoothness and the reported recovery rate represent concrete strengths in an area that often lacks such formalization.","major_comments":[{"comment":"Abstract: the 98.6% recovery rate is reported exclusively on held-out targets 'certified as locally supported.' Certification relies on the same local closeness metric used to retrieve and compose effects, so the result evaluates performance only on the subset where the local smoothness assumption is most likely to hold by selection; this does not test how often the assumption fails or how performance degrades outside the certified set.","section":"Abstract"},{"comment":"Abstract: the error bound for composition is derived under local smoothness, yet no results are provided on the fraction of targets that meet the local-support certification or on performance for non-certified targets; without these, the practical scope of the claimed recovery rate cannot be assessed.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract references human evaluations of bridge experiments and conflict explanations but provides no details on evaluator count, protocol, or agreement metrics; adding these would improve clarity of the supporting evidence.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The available text supplies only the abstract and high-level claims; the absence of detailed methods, data description, or certification procedure verification makes a full technical assessment difficult at present."},"author_rebuttal":{"model":"grok-4.3","summary":"Thank you for the constructive review and for recognizing the potential of the framework along with its formal error bound and recovery results. We address the two major comments on the abstract evaluation below.","responses":[{"response":"The certification step is intentional: it isolates the regime in which the local smoothness assumption (and thus the derived error bound) is empirically supported by the presence of sufficiently close neighbors. Evaluating recovery only on this subset directly tests whether composition succeeds when the modeling assumptions hold, rather than averaging over cases where they do not. This is analogous to reporting in-distribution performance for a method whose guarantees are conditional. We agree that the fraction of targets meeting certification is needed to gauge how often the assumption is satisfied in practice, and we will add this statistic (computed on the held-out set) to the revised abstract and results section. Performance outside the certified set is not claimed to recover effects; non-certification instead triggers the bridge-experiment proposal, which is the appropriate output when local support is absent.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the 98.6% recovery rate is reported exclusively on held-out targets 'certified as locally supported.' Certification relies on the same local closeness metric used to retrieve and compose effects, so the result evaluates performance only on the subset where the local smoothness assumption is most likely to hold by selection; this does not test how often the assumption fails or how performance degrades outside the certified set."},{"response":"We will incorporate the fraction of held-out targets that satisfy the local-support certification criterion into the revised manuscript; this directly addresses the scope question. For non-certified targets the composition step is not executed, so a recovery rate is not applicable; the framework instead outputs a bridge-experiment proposal. The 98.6% figure is therefore presented with the explicit qualifier 'certified as locally supported' to avoid overgeneralization. If additional diagnostics on non-certified cases (e.g., frequency of certification failure) would strengthen the paper, we are prepared to include them.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the error bound for composition is derived under local smoothness, yet no results are provided on the fraction of targets that meet the local-support certification or on performance for non-certified targets; without these, the practical scope of the claimed recovery rate cannot be assessed."}],"tokens_in":1425,"tokens_out":475,"duration_ms":33491,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper introduces ExAtlas to turn scattered social experiments into a linked atlas. It searches for nearby studies in treatment-outcome space, composes their effects under a local smoothness assumption, and then either links the target, reconciles a conflict with moderator suggestions, or proposes a bridge experiment. An error bound is stated for the composition step.\n\nWhat is actually new is the explicit three-outcome workflow built around local composition, plus the attempt to make the archive itself do more of the cumulative work. The human evaluations on bridge plausibility and conflict usefulness give some practical grounding beyond the recovery number.\n\nThe soft spot is the evaluation. The 98.6% direction recovery holds only for held-out targets certified as locally supported. That certification uses the identical closeness criterion that drives the composition, so the test set excludes exactly the cases where neighbors are absent or smoothness fails. The reported figure therefore does not show how often the assumption holds across the archive or how performance drops outside the filtered subset. The abstract does not supply the full certification procedure or the bound derivation.\n\nThis is for computational social scientists who want tools to accumulate experimental results rather than treat each study in isolation. A reader already working on meta-analysis or literature mapping would find the framework idea worth examining.\n\nIt deserves peer review. The mechanism is distinct enough from standard meta-analysis and the outputs are specific enough that referees can usefully check the certification details, the bound, and the human eval protocol.","headline":"ExAtlas claims 98.6% recovery via local effect composition but only on targets pre-filtered by the same closeness metric used for the composition itself.","tokens_in":2363,"tokens_out":379,"would_cite":false,"duration_ms":34256,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"ExAtlas composes effects from nearby studies to link, reconcile or bridge social experiments","keywords":["social experiments","effect composition","conflict reconciliation","bridge experiments","treatment effect surface","local smoothness","atlas mapping"],"falsifier":"A test set of target studies where the composition either fails to recover the observed effect direction at rates much higher than 1.4 percent or where the proposed bridge experiments are judged implausible by domain experts.","tokens_in":2659,"feed_emoji":"🗺️","tokens_out":580,"duration_ms":52725,"temperature":0.7,"pith_summary":"ExAtlas provides a way to build a structured atlas from the thousands of social experiments conducted each year. It locates studies that are close in treatment and outcome space and checks if their effects can be composed to match a new target study. Success with agreement creates links between studies, disagreement triggers reconciliation with possible moderators, and failure leads to suggestions for bridge experiments. This matters because it turns disconnected findings into a map that reveals consistent evidence, explains conflicts, and identifies what is missing, allowing knowledge to accumulate more effectively.","feed_headline":"Atlas maps social experiments via effect composition","feed_subtitle":"Recovers directions in 98.6% of cases and proposes bridges or reconciliations where needed.","key_machinery":"The composition operation on effects from studies close in treatment-outcome space, used to decide linking, reconciling, or bridging.","core_discovery":"The central discovery is that social experiment effects can be composed from locally close prior studies under a local smoothness assumption on the treatment-effect surface. Given a target, the method finds comparable studies, attempts composition, and classifies the outcome as a link if consistent, a conflict to reconcile if inconsistent, or a gap requiring a bridge experiment if composition is not possible. The approach includes an error bound for the composition. It achieves 98.6% recovery of effect direction on held-out locally supported targets, with human raters finding the bridge proposals plausible and the conflict explanations useful for generating theory.","pith_inferences":["The same composition approach might apply to experimental archives in other disciplines such as psychology or medicine.","Conflict explanations could feed into systems that propose and test new theories automatically.","Bridge suggestions could be used to design sequences of experiments that efficiently fill knowledge gaps."],"forward_implications":["Studies with consistent composed effects become linked in the atlas.","Inconsistent compositions lead to proposed moderators or theories to resolve conflicts.","Failed compositions identify gaps and suggest specific bridge experiments.","The experimental archive contains more latent structure than isolated studies reveal."],"fun_headline_variants":["ExAtlas composes effects to link experiment studies","Effect composition builds social experiments atlas","Atlas reconciles conflicts in experiment effects","ExAtlas proposes bridges for gaps in studies","Local composition maps social experiment effects"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Effects change smoothly enough in the local region of treatment and outcome space for composition from nearby studies to work with bounded error.","fun_headline_variants_meta":{"raw":{"variants":["ExAtlas composes effects to link experiment studies","Effect composition builds social experiments atlas","Atlas reconciles conflicts in experiment effects","ExAtlas proposes bridges for gaps in studies","Local composition maps social experiment effects"]},"model":"grok-4.3","cost_usd":0.005472,"raw_usage":{"total_tokens":2579,"prompt_tokens":727,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":54715500,"prompt_tokens_details":{"text_tokens":727,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1800,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":727,"tokens_out":52,"duration_ms":23602,"temperature":1.0,"reasoning_tokens":1800,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T15:58:22.197102+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A test set of target studies where the composition either fails to recover the observed effect direction at rates much higher than 1.4 percent or where the proposed bridge experiments are judged implausible by domain experts.","supporting_citations":[],"review_version":1}