{"id":"6f2a136a-f61d-4e3d-aa73-a4de897351f7","arxiv_id":"2508.03452","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Estimators reconstruct the Curie-Weiss probability measure from partial spin observations, with claimed consistency, asymptotic normality, and large deviation properties.","lead":"This paper proposes statistical estimators that recover the underlying probability measure of a Curie-Weiss model from votes of only a subset of spins. The estimators are claimed to be consistent, asymptotically normal, and computationally cheap, which could help measure social cohesion from survey data.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Identifiability from a fixed small subset is unproven; for a single observed spin the marginal cannot distinguish (β,h), so the 'possibly very small subset' claim needs qualification.","rationale":"The reader's weakest assumption was that the observed subset is representative (exogenous sampling). My concern is different and more fundamental: even under ideal random sampling, the marginal distribution of a fixed subset may not identify (β,h). The abstract explicitly highlights 'possibly very small' subsets, yet for m=1 the marginal is a one-dimensional Bernoulli distribution while the parameter space is two-dimensional. At h=0, the single-spin marginal is exactly 1/2 for all β, so β is not identified. This is not a matter of estimator efficiency; it is a statistical identification failure that no consistent estimator can overcome. If the full paper restricts to m≥2, it must state that clearly, since the abstract's emphasis on very small subsets would otherwise be misleading. My proposed test directly demonstrates the identification failure for m=1. Because the concern is about a missing condition in the central claim, the verdict should be CONDITIONAL: the paper must add and prove the identifiability condition for the smallest subset sizes it claims to handle. I partially agree with the reader's concern about representativeness, but the identifiability issue is a more specific and decisive mathematical risk.","tokens_in":733,"tokens_out":7432,"duration_ms":96717,"concrete_test":"Check the m=1 case explicitly. Compute the single-spin marginal p_{β,h}=P(s_1=1) for the Curie-Weiss model (exactly for small N, or by high-precision simulation for large N) on a grid of (β,h). Verify whether the level sets p_{β,h}=const contain multiple parameter pairs; in particular, confirm p_{β,0}=1/2 for all β (spin-flip symmetry) and find a pair (β_1,h_1)≠(β_2,h_2) with p equal. If such a pair exists, the Fisher information for (β,h) from one observed spin is singular, and the claimed consistency cannot hold for m=1. Also read the full text for any hidden assumption m≥2 or known h; if present, the abstract's 'possibly very small' claim should be amended to 'at least two spins'.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that the Curie-Weiss parameters (β,h) be recoverable from the marginal distribution of a fixed observed subset A. The abstract touts 'possibly very small' subsets, so the extreme case m=|A|=1 must be covered. But for m=1, the data are a single Bernoulli variable with success probability p=P(s_1=1). When h=0, spin-flip symmetry gives p=1/2 for every β, so β is not identifiable from one spin. More generally, the map (β,h)↦p is a single real number from a two-dimensional parameter space, so it cannot be injective; varying β and h along a level curve leaves p unchanged. Therefore no estimator based on one observed spin can be consistent for the full probability measure. If the paper's consistency results require m≥2 or require h known, that restriction is absent from the abstract, and the advertising of 'very small' subsets is misleading. The full text must state the asymptotic regime (fixed m vs m→∞) and prove identifiability of (β,h) from P_{β,h}^A; without such a proof, the central claim fails in the small-m regime it advertises.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript, as presented by its abstract, studies estimation of the Curie-Weiss model parameters from observations of spins in a subset of the population. It claims that the proposed estimators are consistent, asymptotically normal, and satisfy large deviation principles, while requiring only a (possibly very small) subset of spin realisations and being computationally cheap.","tokens_in":922,"tokens_out":5627,"duration_ms":65602,"significance":"If the claims are established under appropriate conditions, the paper would offer a practical and low-cost method for estimating social cohesion from partial survey data, with relevance to political science and social choice. However, the abstract alone does not provide the needed identifiability conditions and proof structure, and the 'very small subset' claim is in direct tension with non-identifiability when only one spin is observed. The paper's significance therefore hinges on the full text supplying precise conditions and rigorous proofs.","major_comments":[{"comment":"The claim that the estimators work with a 'possibly very small' subset is unqualified. For m=|A|=1, the observed data are a single Bernoulli random variable with probability p=P(s_1=1); when h=0, spin-flip symmetry forces p=1/2 for every β, so β is not identifiable from the marginal. More generally, the map (β,h)↦p cannot be injective from R^2 to [0,1]. Hence no estimator based on a single spin can consistently reconstruct the full probability measure. The manuscript must either restrict the main results to m≥2 or m growing, and prove identifiability from P^A_{β,h}, or state additional assumptions that make m=1 identifiable.","section":"Abstract"},{"comment":"The abstract does not specify the sampling mechanism for the observed subset. If the subset is not independent of the spin values, the observed marginal distribution differs from P^A_{β,h}, and consistency and asymptotic normality cannot be expected in the stated form. The full text must define 'representative subset' precisely and include results under that sampling scheme.","section":"Abstract"},{"comment":"The properties 'consistency, asymptotic normality, and large deviation principles' are asserted without specifying the estimators, the parameter space, or the underlying regularity conditions. The full text must contain precise theorem statements and proofs; in particular, a large deviation principle requires exponential tightness and identification of the rate function, which cannot be inferred from the abstract.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'some positive properties' is too vague; replace it with a precise enumeration of the proven results.","section":"Abstract"},{"comment":"The abstract uses 'reconstructing the probability measure' and 'estimators' without clarifying whether the target is the full measure or the parameters (β,h); please align the terminology.","section":"Abstract"},{"comment":"The term 'representative subset' is a term of art in survey sampling; give a mathematical definition or a reference to the sampling scheme used.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript as submitted for review is an abstract-only document, which makes a standard referee assessment impossible. The identifiability objection at m=1 is mathematically definite and requires an explicit qualification or proof. I recommend that the editor require a full manuscript with rigorous statements and proofs before considering the paper further."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know quickly. First, the problem is real and the abstract makes concrete promises: consistent estimators of the Curie-Weiss measure from partial spin observations, with asymptotic normality and LDPs, at low computational cost. That is a useful target, especially for social choice and survey contexts. Second, the identifiability claim is not secured as stated. The stress-test note is correct: with a single observed spin (m=1), the marginal is a single Bernoulli probability, so (β,h) is not identifiable when h=0 (spin-flip symmetry) or generally along level curves. If the paper's theorems require m≥2, known h, or an asymptotic regime where m grows, the abstract's 'possibly very small subset' is misleading and the full text must say so. This is not a manufactured flaw; it is a load-bearing condition.\n\nWhat the paper does well, based on the abstract, is pick a practical inference problem and commit to specific asymptotic guarantees. That is more than most abstracts do. I cannot verify novelty or soundness without the full text, and the reference list is absent. But the stated properties are the right ones to aim for, and the mean-field structure of Curie-Weiss makes an exact analysis plausible.\n\nSoft spots, in proportion: (1) the identifiability gap above; (2) no comparison to existing estimators for Curie-Weiss or to plug-in methods using only observed marginals; (3) no statement of the sampling model for the observed subset beyond 'representative', which matters if selection depends on spin values. These are not damning in themselves, but they must be addressed.\n\nWho is this for? Statisticians and probabilists working on mean-field inference, and applied people in political science or sociology who need interpretable measures of social cohesion. If the theorems hold under explicitly stated conditions, the paper will be a solid contribution to a niche literature.\n\nRecommendation: send it to a serious referee. The topic is meaningful, the claims are specific, and the identifiability condition is checkable. Desk rejection would be wrong; acceptance without fixing the small-subset qualification would also be wrong. I would want the referee to demand a precise identifiability theorem.","headline":"Useful estimator for Curie-Weiss from partial observations, but identifiability from a single spin is unproven and needs explicit conditions.","tokens_in":1413,"tokens_out":2252,"would_cite":false,"duration_ms":25479,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60F05","60F10","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"The probability measure of a Curie-Weiss model can be reconstructed consistently and with asymptotic normality from observations of only a subset of the spins.","keywords":["Curie-Weiss model","partial observations","probability measure reconstruction","consistency","asymptotic normality","large deviations","social cohesion","survey sampling"],"falsifier":"Generate Curie-Weiss data with known parameters for a large population, then keep only spins whose values are above the median of the full hidden configuration; apply the paper's estimator to this dependent subsample and check whether the estimated measure converges to the true one as the population size increases.","tokens_in":541,"feed_emoji":"🗳️","tokens_out":4904,"duration_ms":52496,"temperature":0.7,"pith_summary":"This paper asks whether the probability measure of a Curie-Weiss model—a mean-field model of binary decisions in which every agent interacts equally with every other—can be reconstructed from observing the 'votes' of only a subset of the population. The authors study estimators that use the observed subset and show that these estimators are consistent, asymptotically normal, and obey large deviation principles, while remaining computationally cheap. If correct, this means social cohesion, measured as the interaction strength in the model, can be estimated from small representative surveys rather than from full population data. The potential payoff is broad: political science, sociology, automated voting, and preference aggregation all use such mean-field models.","feed_headline":"Partial voter samples can recover the full Curie-Weiss measure","feed_subtitle":"Small representative survey subsets yield consistent, asymptotically normal estimates of group influence.","key_machinery":"The central object is the Curie-Weiss probability measure on $N$ binary spins, a mean-field model where the probability of a configuration depends only on the total sum of the spins through a coupling parameter. The machinery is the class of estimators constructed from the empirical distribution of the observed subset, together with the probabilistic tools (laws of large numbers, central limit behaviour, and large deviation estimates) that give the estimators their consistency, asymptotic normality, and exponential tail bounds.","core_discovery":"The central claim is that partial observation does not obstruct statistical recovery of the underlying Curie-Weiss measure. Specifically, the paper asserts that from the realised votes of a possibly very small subset of the population one can build estimators of the model's probability measure that converge to the true measure as the population grows, are asymptotically normal, and satisfy large deviation principles. The reconstruction is not a physical measurement but a statistical estimation problem, and the paper's contribution is to show that the interaction structure that generated the votes can be recovered with these guarantees from a subset alone.","pith_inferences":["The abstract does not state it, but the representativeness of the subset is likely a missing-at-random assumption; if the sampling rule depends on the hidden spin values, the estimators would probably be biased.","A natural transfer is to other mean-field exponential-family models where only a subset of coordinates is observed; the same estimator logic should apply.","For survey practice, the result suggests that random subsampling of respondents is the safe design that satisfies the paper's conditions."],"forward_implications":["If the estimators work, social cohesion can be measured from partial survey data covering only a small fraction of a population.","Consistency means the reconstructed measure improves as the full population grows, even when the observed subset stays relatively small.","Asymptotic normality gives approximate confidence intervals for the reconstructed interaction strength.","Large deviation principles provide exponential tail bounds, useful for evaluating the reliability of the estimate.","Low computational cost makes the method practical for large-scale data in political science, sociology, and automated voting."],"supporting_citations":[],"fun_headline_variants":["Subset votes reveal full Curie-Weiss measure","Small voter subset recovers group influence","Partial samples reconstruct Curie-Weiss measure","Curie-Weiss recoverable from partial votes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the observed subset of the population is representative, meaning the unobserved spin values do not influence which spins are sampled; if sampling depends on the hidden votes, the consistency and normality guarantees can fail.","fun_headline_variants_meta":{"raw":{"variants":["Subset votes reveal full Curie-Weiss measure","Small voter subset recovers group influence","Partial samples reconstruct Curie-Weiss measure","Curie-Weiss recoverable from partial votes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000562,"raw_usage":{"total_tokens":2618,"prompt_tokens":848,"completion_tokens":1770,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":1713}},"tokens_in":464,"tokens_out":1770,"duration_ms":12792,"temperature":1.0,"reasoning_tokens":1713,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:25:07.608513+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate Curie-Weiss data with known parameters for a large population, then keep only spins whose values are above the median of the full hidden configuration; apply the paper's estimator to this dependent subsample and check whether the estimated measure converges to the true one as the population size increases.","supporting_citations":[],"review_version":1}