{"id":"456f81ed-7e95-4e9f-99ea-e2e966350fd1","arxiv_id":"2608.06968","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"State encodings, not just the physical state, determine the collective synchronization outcomes of language-model agent populations, and the effect is model-family dependent.","lead":"Three ways of describing the same circular phase state, as circular moments or as two differently labelled histograms, made populations of language-model agents synchronize or fail to synchronize on identical physics. The result survives a matched replay control in three model families, with the encoding-to-outcome map reversing between GPT and Claude.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proprietary API stationarity is the weakest load-bearing assumption; the controls make provider drift an unlikely confound, so the verdict stands.","rationale":"The paper is a carefully controlled empirical demonstration. The central claim rests on two pillars: (1) collective experiments in GPT and Claude showing encoding-dependent synchronization outcomes, and (2) identical-field replay showing operator-level effects in three model families. Both pillars use appropriate inference units (physical seeds and physical fields, respectively), prespecified thresholds, exact permutation tests, and comprehensive negative controls (K=0, field-blocked label permutation, block-noise floors). The cross-family reversal is treated correctly as evidence for model-dependence rather than a failed replication. The weakest link is indeed the stationarity of proprietary API responses, exactly as the reader identified. The absence of an echoed temperature for macroscopic runners and the finite acquisition windows leave a small opening for provider-side drift. However, the internal controls (K=0 exact coincidence, small block noise, fixed run order, and the magnitude of the effects) make this an unlikely confound, and the replication in two independent model families with opposite directions further reduces the plausibility of a time-dependent artifact. No internal inconsistency or circularity was found. A concrete test of API stationarity would strengthen confidence, but the current evidence supports the ACCEPT verdict. The concern does not change the verdict because it is a reasonable environmental assumption given the locked model IDs and short acquisition windows, not a demonstrated flaw.","tokens_in":57957,"tokens_out":6407,"duration_ms":71820,"concrete_test":"Re-run the GPT matched collective experiment (moments versus centers, K=0.08, six physical seeds) with temperature explicitly set (e.g., 0.7 or 1.0) and with run order randomized across encodings within each seed; then repeat the same acquisition two weeks later under the same locked model IDs. If the moments 6/6 versus centers 0/6 locking separation reproduces under both random-order conditions and at both time points, the stationarity assumption is validated. As a stronger, provider-independent check, replicate the core collective comparison with an open-weight model (e.g., Llama-3.1-70B-Instruct) at a fixed checkpoint with temperature=0 and deterministic sampling, which eliminates provider-side variability entirely.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that state encodings form part of a model-dependent effective interaction law requires that the observed encoding effects are not artifacts of non-stationarity in the proprietary models' response distributions over the acquisition period. The paper locks model IDs and most generation parameters, but for the macroscopic collective runners temperature is not echoed in the resolved configuration (Supplementary Table S2), and provider-side updates or sampling drift during the 1.6 h GPT acquisition could in principle create encoding-correlated shifts. The run order was fixed to avoid wall-clock association, but a fixed order cannot fully rule out non-monotonic drift or an interaction between time and the specific task sequence. If such drift occurred, the 6/6 versus 0/6 locking separation could be inflated. Internal controls weaken this possibility: the exact K=0 control shows identical trajectories despite differing actions, and the identical-field replay block noise is small (0.092 versus 0.344 separation), so large within-window drift would have been visible. The Claude reversal, with separation ratios around 20x, would require a coincidental time-dependent reversal across families. This is an environmental and tooling assumption, not an internal logical flaw, and it does not by itself undermine the claim as scoped.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a circular-synchronization experiment in which pretrained language-model agents observe only an encoded summary of their neighbours' relative phases and choose to advance, stay, or retard. The authors manipulate the state encoding (first three circular moments vs 24-bin histograms labelled by bin centers or bin intervals) while holding the physical field fixed. In GPT, the moments encoding synchronizes the population in 6/6 seeds at two positive couplings, whereas both histogram encodings lock in 0/6; in Claude, the encoding effect replicates but reverses, with the histogram encodings locking. Identical-field replay on 48 frozen fields shows that the presented encoding alone changes the advance/stay/retard action distribution in GPT, Claude, and Gemini. Additional controls show that presentation changes with moment values fixed shift the GPT operator. The authors conclude that state encodings form part of a model-dependent effective interaction law, not a neutral interface.","tokens_in":58128,"tokens_out":10230,"duration_ms":105392,"significance":"If the result holds, it has direct implications for the design and evaluation of multi-agent LLM systems, for performative prediction, and for the interpretation of state-representation choices in agent scaffolds. The study is methodologically careful: matched physical seeds across encodings, an exact K=0 negative control, identical-field replay with field-blocked label permutations, prespecified thresholds and sensitivity sweeps, and full release of raw traces and code. The cross-family reversal is a strong falsification of any claim that one encoding is intrinsically superior. The main limitation is reliance on proprietary API behaviour, which the internal controls mitigate but cannot fully eliminate.","major_comments":[],"minor_comments":[{"comment":"In the first paragraph, the string 'reserveserializationfor' should read 'reserve serialization for'; the missing spaces make the sentence difficult to parse on first reading.","section":"Introduction"},{"comment":"The temperature is not echoed in the resolved configuration for the macroscopic collective runners; because the stationarity of the proprietary API is the main external assumption, consider reporting the effective temperature used or adding a time-indexed stability analysis of action distributions over each acquisition window (e.g., first versus second half of calls) to further rule out drift as a confound.","section":"Supplementary Table S2"},{"comment":"The reformatting contrast (original–reformatted mean 0.145, 95% CI 0.079–0.219) overlaps the test–retest floor (0.049–0.090), and the text correctly notes this; consider adding a similar explicit qualifier in the abstract so that 'presentation alone shifted the operator' is not misread as being driven by the weak reformatting effect, when the padding condition carries the effect.","section":"Fig. 5 caption"}],"recommendation":"minor_revision","confidential_remarks":"I concur with the stress-test assessment: the proprietary-API stationarity concern is mitigated by the internal controls, and the scientific claim is well supported. The manuscript is within scope for physics.soc-ph. I recommend minor revision to address the presentation issues listed above."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the take: this paper earns a serious referee. It goes beyond the known single-turn prompt sensitivity by showing the effect survives closed-loop feedback and flips collective synchronization outcomes in matched seeds, with a clean reversal between GPT and Claude. The design is unusually careful: matched physical seeds, an exact K=0 control where trajectories coincide bit-identically despite different actions, an identical-field replay with field-blocked permutations, prespecified thresholds, and full release of raw traces. That covers the usual ways these results go wrong.\n\nWhat's actually new: state encodings, not just physical states, become part of the effective interaction law in an agent population. The claim is scoped properly—it's about model-encoding pairs, not a universal ranking. The cross-family reversal is treated as evidence for model dependence, not as a failed replication, which is right.\n\nThe soft spots are minor and mostly acknowledged. The proprietary model stability assumption is the weakest link: temperature isn't echoed in the resolved config for the macro runners, and provider drift over the acquisition window could in principle be encoding-correlated. But the internal controls make that unlikely—the K=0 control and the small block noise would catch large drift, and the Claude reversal would need a coincidental time-dependent reversal. N=17 is small and finite-size; macro results cover only two model families; the models are non-frontier. None of this undermines the claim as scoped.\n\nThe circularity burden is low. The Fourier coefficients describe measured response curves but don't construct the synchronization result; no fitted parameter defines the outcome. The K=0 and replay controls break any loop where the encoding merely labels what the model was trained to say.\n\nWho gets value: anyone building or evaluating multi-agent LLM systems, especially those reporting benchmarks without versioning the observation serializer. Also people studying performative prediction and effective policies. I'd cite this in my own work if I were writing about multi-agent evaluation.\n\nRecommendation: accept for peer review. It's not a paradigm shift, but it's a solid empirical result with real methodological care, and the claim is exactly as large as the evidence.","headline":"Careful, well-controlled demonstration that observation encodings change collective LLM-agent dynamics; the API-stationarity worry is real but the controls contain it.","tokens_in":58693,"tokens_out":2532,"would_cite":true,"duration_ms":24821,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["05.45.Xt","89.65.-f"],"model":"deepseek-v4-flash","headline":"How a state is described to a language-model agent is part of its effective interaction law, not a neutral interface: moment encodings synchronized GPT populations 6/6 seeds and histogram encodings 0/6, with the effect reversed in Claude.","keywords":["language-model agents","state encoding","observation map","synchronization","collective dynamics","phase oscillators","performative feedback","multi-agent systems"],"falsifier":"A prospective replication that pins provider model versions and records all sampling parameters, reruns the identical-field replay on fresh fields, and finds the between-encoding total-variation distance fallen to the within-encoding test-retest floor would falsify the claim that the encoding itself moves the operator. A second decisive test is the operator-swap the paper leaves open: if the closed-loop trajectories were unchanged when the action distribution of one encoding is substituted for another on a fixed state sequence, the measured microscopic differences could not be what selects the collective outcome.","tokens_in":57702,"feed_emoji":"🤖","tokens_out":13402,"duration_ms":119922,"temperature":0.7,"pith_summary":"Coupled language-model agents never act on an environment directly; they act on a text description of it, and this paper argues that the description — the observation map — is part of the agents' effective interaction law rather than a neutral interface. To test this it built a minimal circular-synchronization system in which each agent sees only an encoding of its neighbors' relative phases and chooses to advance, stay, or retard, while a deterministic engine applies the coupling. In GPT, encoding the same physical field as low-order circular moments locked the population in 6 of 6 seeds at positive coupling, while two histogram encodings locked in 0 of 6; in Claude the encoding effect replicated but reversed direction. On 48 frozen fields replayed from real runs, re-encoding alone moved each agent's action probabilities by a mean total-variation distance of 0.344, against a 0.092 test-retest floor, an effect seen in all three model families tested. A sympathetic reader takes the conclusion to be that the observation serializer is a model-dependent component of collective dynamics, so evaluations of LLM agent populations must report and re-validate it together with the model.","feed_headline":"State description alone flipped whether AI agents synchronized","feed_subtitle":"The same physical state, described differently, picked opposite synchronization outcomes in GPT and Claude.","key_machinery":"The load-bearing object is the circular-synchronization assay, a population of $N = 17$ phase oscillators whose only coupling is the sampled action of a pretrained language model: the engine applies $x_i(t+1) = x_i(t) + \\omega_i + K f_i(t)$ with $f_i(t) \\in \\{-1, 0, +1\\}$, and the model sees only a text serialization of the relative-phase field. The companion object is the response operator $p_R(\\rho)$, the advance/stay/retard probability distribution measured from repeated calls at a fixed field $\\rho$ under observation map $R$; the three maps are the first three circular moments, a 24-bin histogram labelled by bin centers, and the same 24 masses labelled by bin intervals. The decisive mechanism is identical-field replay, which takes 48 frozen fields generated by the closed loop and re-encodes each under all three maps, so any response difference can only be attributed to the encoding. Supporting machinery includes the order parameters $r_m(t)$, the Fourier decomposition of the signed response curve as a data-driven phase interaction function, and the $K = 0$ control arm in which the coupling term vanishes, making trajectories coincide exactly while actions still differ.","core_discovery":"The paper's central claim is that the observation map, the transformation from physical state to model input, is a constituent of the effective policy of a language-model agent, so that different encodings of the same state can drive a population to qualitatively different collective outcomes. The demonstration proceeds in three layers. On identical replayed fields, the three encodings — moments, bin centers, bin intervals — produced action distributions separated by a mean pairwise total-variation distance of 0.344, against a between-block test-retest floor of 0.092, with no separation this large in 5,000 field-blocked relabellings ($p = 0.0002$), and the two information-matched histogram serializations alone separated by 3.2 times the floor. In the closed loop, this microscopic difference redirected collective outcomes: in GPT the moments encoding locked in 6/6 seeds at both positive couplings while the histogram encodings locked in 0/6 (exact paired sign test $p = 0.03125$); in Claude the same design produced the reverse ordering, with moments locking 0/6 and intervals 6/6 (exact within-seed enumeration $p = 0.00103$ for terminal lock). The exact $K = 0$ control, in which the coupling term vanishes, gave identical trajectories across encodings while actions still differed, excluding matched-engine and initial-condition confounds. The stated conclusion is that state encodings form part of a model-dependent effective interaction law, not a neutral interface.","pith_inferences":["An untested corollary of the model-dependence result is that the same encoding-to-outcome map should be mutable by training: fine-tuning or instruction-tuning a single model family would be predicted to shift, and possibly reverse, which encoding synchronizes, offering a cheaper test of the mechanism than adding new model families.","A quantitative gap the authors flag is whether the measured operator differences are sufficient to reproduce the encoding-specific trajectories; an operator-swap experiment at a fixed state sequence would close it and would also separate the contribution of the operator from that of trajectory feedback.","If the phenomenon generalizes beyond this minimal assay, then design choices such as memory, message-passing interfaces, and structured observation schemas in deployed multi-agent systems acquire the same status as the reward or instruction contract, and neutral-interface assumptions in agent scaffolding should be abandoned.","The sharp, grid-limited activation of the moments operator near zero directional imbalance hints at a possible threshold-like sensitivity to peer asymmetry; testing finer imbalance grids would reveal whether the response is continuous or genuinely discontinuous."],"forward_implications":["The observation serializer becomes a versioned component of any LLM agent evaluation: a result obtained with one encoding is a property of the encoding–model pair, not of the task.","Cross-model comparisons of agent populations cannot be transferred by analogy, because the encoding-to-outcome map reversed between GPT and Claude under an identical design.","Single-turn prompt-sensitivity measurements understate the effect of formatting on collective dynamics, because feedback can amplify, suppress, or invert a small operator difference.","Surrogate models and simulators that stand in for LLM agents must be validated on the field distribution the interacting population actually generates, not only on the acquisition distribution.","Because centers and intervals carried identical bin masses and still separated the operator, information-equivalent serializations are not behaviorally equivalent, so a faithful encoding cannot be defined by information content alone."],"supporting_citations":[{"why":"Demonstrates that LLM outputs are sensitive to prompt formatting, the single-turn effect the paper extends to the feedback setting.","marker":"[1]"},{"why":"Shows output sensitivity to the order of information in prompts, one of the formatting sensitivities the design holds fixed while varying the observation map.","marker":"[2]"},{"why":"Defines performative prediction, the feedback loop in which a deployed policy changes the data it is later evaluated on, which frames the closed-loop argument.","marker":"[14]"},{"why":"States the classical result that policies act on observations rather than latent states, grounding the claim that the observation model is part of the effective policy.","marker":"[17]"},{"why":"Introduces the coupled-phase-oscillator framework from which the circular synchronization assay is drawn.","marker":"[18]"},{"why":"Provides the order parameters $r_m(t)$ and synchronization theory used to classify the collective outcomes.","marker":"[20]"},{"why":"Defines the phase order (interaction) function that the measured signed response curve $g_R(\\delta)$ is the data-driven analogue of.","marker":"[26]"},{"why":"Shows phase interaction functions can be recovered from measured oscillator responses, the template for reconstructing $g_R(\\delta)$ from language-model queries.","marker":"[27]"},{"why":"Argues that multi-agent LLM evaluations are unreliable and validation should not transfer across setups, supporting the paper's cross-family caution.","marker":"[29]"}],"fun_headline_variants":["State encoding flips AI agent synchronization outcomes","Same state, different encoding: AI collectives sync differently","Encoding alone selects AI population sync in GPT and Claude","How you describe state changes AI agents' collective dynamics","Model-dependent encoding governs AI synchronization, not state"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each proprietary model's response behavior is stationary over the acquisition windows, so that differences between encodings reflect the encoding rather than drift in the model or its provider backend; the paper locks model identities and most generation parameters, but temperature is not echoed for the macroscopic collective runners and provider-side updates during acquisition cannot be fully excluded.","fun_headline_variants_meta":{"raw":{"variants":["State encoding flips AI agent synchronization outcomes","Same state, different encoding: AI collectives sync differently","Encoding alone selects AI population sync in GPT and Claude","How you describe state changes AI agents' collective dynamics","Model-dependent encoding governs AI synchronization, not state"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000855,"raw_usage":{"total_tokens":3748,"prompt_tokens":1014,"completion_tokens":2734,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":2660}},"tokens_in":630,"tokens_out":2734,"duration_ms":22411,"temperature":1.0,"reasoning_tokens":2660,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:27:40.467651+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective replication that pins provider model versions and records all sampling parameters, reruns the identical-field replay on fresh fields, and finds the between-encoding total-variation distance fallen to the within-encoding test-retest floor would falsify the claim that the encoding itself moves the operator. A second decisive test is the operator-swap the paper leaves open: if the closed-loop trajectories were unchanged when the action distribution of one encoding is substituted for another on a fixed state sequence, the measured microscopic differences could not be what selects the collective outcome.","supporting_citations":[{"cited_title":"Z., Zhai, Y","cited_arxiv_id":null,"evidence_quote":"Shows phase interaction functions can be recovered from measured oscillator responses, the template for reconstructing $g_R(\\delta)$ from language-model queries."}],"review_version":1}