{"id":"8529b9ac-6704-4e8e-ad69-1b82f26c18c5","arxiv_id":"2508.15408","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In grouped panel data, tiny groups are hard to estimate and standard information criteria can pick the wrong number of groups; this paper derives when estimation works and proposes modified criteria.","lead":"This paper studies what happens when some groups in a panel dataset are very small, and shows that estimating those groups only works if the time dimension is large enough. It also proposes a new rule for choosing the number of groups that handles small groups better than existing rules.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MIC's improvement over existing criteria is unproven; Remark 3 admits no consistency result, and the finite-sample evidence is a 100-replication simulation with an arbitrary tuning constant.","rationale":"The reader correctly identified Assumption 2(g) and the simulation evidence as weak points, but the most load-bearing issue is the omitted proof of consistency for the proposed MIC. The rate condition is a stated sufficient condition and is internally consistent with the proofs; the paper candidly flags it as a limitation. In contrast, the MIC's improvement over existing criteria is a central positive claim, and Remark 3 explicitly concedes that its consistency is not proved. This means the headline contribution is supported only by a small, potentially in-sample simulation with an arbitrary constant. This reinforces the reader's CONDITIONAL verdict rather than overturning it: the theoretical parts of the paper are valuable, but the MIC's practical superiority is not established with the same rigor. The proposed concrete test would either confirm the simulation evidence or reveal that the MIC's performance is fragile.","tokens_in":42189,"tokens_out":20916,"duration_ms":224964,"concrete_test":"Re-run the Monte Carlo for DGP1 (static panel) with N=120, T=10, α=0.2, using 10,000 replications and three variants of hMIC1 with scaling constants c ∈ {0.25, 0.5, 1} for N>T. Report the selection frequency of the true K0=3 for MIC, Bai-Ng, and BIC, with standard errors. If the true selection frequency of MIC is not significantly above that of Bai-Ng/BIC, or if it drops below 0.9 for any c, then the finite-sample superiority claim is not robust to the arbitrary constant and does not support the abstract's claim that MIC performs well.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's third contribution is a modified information criterion (MIC) 'designed to perform well' in the presence of small groups. The central claim that MIC improves on Bai-Ng and BIC penalties is not theoretically established. Remark 3 (Section 3.4) explicitly states: 'we do not prove the consistency of K-hat(h_MIC1_NT).' The design rules (i)-(iv) in §3.4 only ensure that MIC avoids the specific under-estimation condition (5); this is necessary but not sufficient for consistent selection. No theorem shows that K-hat(h_MIC1) converges to K0 under any data-generating process. The claimed superiority rests on a Monte Carlo study with only 100 replications, no reported standard errors, no code or data, and an arbitrary scaling constant (0.5 for N>T) whose sensitivity is explicitly left unexplored (Remark 8 in §3.4). Moreover, the simulation design itself motivated the penalty choice, so the evidence is partly in-sample. This is load-bearing because the abstract's headline promise—that the proposed MIC 'allows one to discover small groups without producing too many groups'—would be vacuous if MIC is inconsistent or if its good performance depends critically on the 0.5 constant. The paper's own limitation statement in Remark 3 confirms this gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies K-means / least-squares clustering in panel data models when some latent groups are 'small,' i.e., of size N^{α_k} with α_k < 1. It makes three contributions. First, it derives sufficient conditions under which the LS slope and grouped fixed-effects estimators are consistent and asymptotically normal (Theorems 1-2), with the key rate condition that T must grow fast enough relative to N and the smallest group size. Second, it derives sufficient conditions for information criteria to under-estimate the number of groups (Propositions 1-2) and shows that the Bai-Ng penalty can satisfy these conditions, while the BIC penalty over-estimates in finite samples in models without grouped fixed effects. Third, it proposes modified information criteria (h_MIC1, h_MIC2) designed to avoid the identified under-estimation condition. The paper reports a 100-replication Monte Carlo study and an empirical application to Japanese firms' sales growth. The paper explicitly states in Remark 3 that no consistency result is proved for the proposed MIC.","tokens_in":42517,"tokens_out":17849,"duration_ms":205263,"significance":"The theoretical parts of the paper, if correct, fill a real gap: most clustered panel asymptotic theory assumes all group sizes are proportional to N. The paper shows the consequences of relaxing this assumption, and the proofs in the appendix are detailed and self-contained. The characterization of when Bai-Ng and BIC-type penalties fail is also of practical value. However, the proposed MIC is the least theoretically supported contribution: the design rules only ensure that a particular sufficient condition for under-estimation is avoided, and the finite-sample evidence is based on 100 replications without standard errors or sensitivity analysis. The paper is therefore more significant for its theoretical characterization of small-group asymptotics than for the specific new criterion as currently justified.","major_comments":[{"comment":"The paper's third contribution is the MIC, and the abstract credits it with allowing one to 'discover small groups without producing too many groups.' Yet no consistency or selection-rate theorem is proved for h_MIC1 or h_MIC2. The design rules (i)-(iv) ensure that condition (4) holds and that the sufficient under-estimation condition (5) fails, but failure of a sufficient condition is not a proof of correct selection; other channels of inconsistency remain possible. Remark 3 concedes this. This is load-bearing: if the only formal property is 'not known to under-estimate through one specific mechanism,' the abstract's promise is not backed by theory. I would ask for either a consistency result under additional conditions (e.g., a lower bound on the population gap between the true and under-fitted models), or a clear reframing of the MIC as a heuristic accompanied by a substantially stron","section":"§3.4, Remark 3"},{"comment":"The finite-sample evidence for the MIC is thin. The Monte Carlo uses only 100 replications and reports only means, with no standard errors, confidence bands, or full distributions. The penalty constant 0.5 in h_MIC1 is chosen in part because of the simulation results, and footnote 8 explicitly says the sensitivity of the scaling constant is not explored. The design of the simulation also motivated the penalty choice, so the evidence is partly in-sample. Please report Monte Carlo standard errors (or interquartile ranges), increase the number of replications, and provide a sensitivity analysis over the scaling constant (e.g., 0.25, 0.5, 1, 2) and over T/N ratios. Without this, the claim of 'good performance' is not robustly documented.","section":"§5.1, Figures 1-3, footnote 8"}],"minor_comments":[{"comment":"The text says that for DGPs 1 and 3 a within-transformation is applied to demonstrate robustness to individual fixed effects. DGP 3 is a grouped fixed-effects model, not an individual fixed-effects model. Please clarify exactly what transformation is applied and why it is compatible with the theoretical assumptions, since within-transformation does not eliminate group-time effects.","section":"§5, paragraph before §5.1"},{"comment":"Group 3 in the empirical application contains only 3 firms. The reported significance stars for the slope estimates in this group should be interpreted with great caution; a sentence acknowledging the extremely small group size as a limitation would be appropriate.","section":"Table 3"},{"comment":"The sentence 'we do not explore in this direction' should be moved or expanded in the main text, because the arbitrary scaling constant is a central concern for a practitioner using the MIC.","section":"§3.4, footnote 8"},{"comment":"The paper would benefit from a data and code availability statement, especially because the simulation study is used to support the main practical recommendation.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The theoretical results (Theorems 1-2 and Propositions 1-2) appear sound and are a useful extension of Bonhomme and Manresa (2015). The main risk to the paper is the MIC contribution: as written, the abstract and introduction promise more than the theory delivers, and the simulation evidence is not strong enough to carry that weight. If the authors can either add a formal selection-consistency result or substantially reframe the MIC as a purely practical heuristic and strengthen the Monte Carlo evidence, I would be supportive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take.\n\nThe new stuff that matters: the paper relaxes the proportional-group-size assumption in LS/K-means panel clustering. Theorem 1 gives sufficient conditions for consistency and asymptotic normality when some groups shrink like N^alpha, and makes precise that T must grow faster as groups get smaller. The GFE extension in Theorem 2 is similar. The inconsistency results in Propositions 1 and 2 are genuinely new: they characterize when Bai-Ng and BIC-type penalties under- or over-estimate the number of groups. Those proofs check out at a glance; Appendix A and B are detailed, the rate conditions cohere, and the derivations are self-contained. This is a real contribution to the panel clustering subfield.\n\nWhat I'm less sold on is the modified information criterion. The paper itself admits in Remark 3 that it does not prove consistency of the MIC, and the design rules only avoid one specific underestimation condition. That's necessary, not sufficient. So the headline promise about MIC \"allowing one to discover small groups\" rests on a 100-replication simulation with no error bars, no code or data, and a tuned 0.5 scaling constant. That evidence would be fine as a heuristic demonstration, but it is not an established claim. The empirical application illustrates the behavior rather than validates the method. The stress-test note gets this right.\n\nI also think Assumption 2(g) is a genuine limitation: it requires T to grow faster than N in the worst case, and the practical usefulness in short panels is limited. But the paper flags this clearly, and it is an honest statement about the estimator's behavior rather than a hidden defect.\n\nWho should read it: econometricians working on grouped panel data. It deserves a serious referee. The theoretical part deserves to be in the literature, and the MIC can be presented for what it is—an adjustment worth investigating. My recommendation: send to a good econometrics journal, but the referee should push for either a consistency proof for the MIC (whether positive or negative) or a much more careful simulation with more replications, error bars, and sensitivity analysis for the tuning constant. If the author can't supply proof, they should tone down the claims.","headline":"Solid theoretical extension of Bonhomme–Manresa to small groups; the MIC part is a heuristic with no consistency proof and thin simulations, but the inconsistency results and rate conditions are the real contribution.","tokens_in":42932,"tokens_out":1499,"would_cite":true,"duration_ms":17232,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows K-means panel clustering can recover tiny groups, provided the time dimension grows fast enough; standard model-count rules fail in this setting, and new MIC rules fix them.","keywords":["panel data","group structure","small groups","K-means clustering","information criterion","grouped fixed effects","asymptotic normality","model selection"],"falsifier":"Simulate a panel with N = 200, T = 10, a small group of size about N^{0.3} (roughly 5 units), well-separated slope parameters, and group sizes otherwise proportional to N. If the proposed MIC selects more than the true K0 groups, or if the K-means membership error rate fails to vanish as N grows with T fixed, the central claim that T must grow like N^{1-2 alpha} would be contradicted. More directly, run the same DGP with T = c N^{0.8} versus T = c N^{1.2} for alpha = 0.3; the theory predicts the first fails to classify the small group perfectly in large samples, while the second succeeds.","tokens_in":2003,"feed_emoji":"📊","tokens_out":3912,"duration_ms":88288,"temperature":0.7,"pith_summary":"Relaxing the usual assumption that every latent group is proportional to N, this paper considers panels where some groups grow like N^alpha with alpha < 1. It proves the least-squares (K-means) estimators of slopes, grouped fixed effects, and memberships remain consistent and asymptotically normal only if the time dimension T grows fast enough relative to the smallest group's size. It also shows that standard information criteria for the number of groups—Bai and Ng's and the usual BIC—can be inconsistent or unreliable when small groups are present. The proposed modified information criteria (MIC) avoid both missing small groups and over-splitting. This matters because real panel applications often produce small groups, and the paper's empirical example finds a three-firm group that standard criteria miss.","feed_headline":"Small groups in panels need longer samples for K-means","feed_subtitle":"Standard information criteria miss rare subgroups; the paper's new penalty finds them without over-splitting.","key_machinery":"The central object is the group-size exponent alpha_k, with N_k = tau_k N^{alpha_k}, and the constructed partition Gamma_N(K,m) from Definition 1. For any K < K0, Gamma_N(K,m) keeps the large groups nearly intact and pools the small groups; Lemma A.4 shows its least-squares fit differs from the true K0-group fit by only Op(N^{alpha_{m+1}-1}). This vanishing gap is why standard penalties fail: they cannot detect a misspecified model whose fit is already essentially optimal. In the grouped-fixed-effects model, the same construction uses the matrix D_{T,i}, and the effective penalty becomes T h_{NT}, which drives the different proposed penalty scale.","core_discovery":"The central claim is that small groups are not a nuisance: they change the rate conditions for inference. Writing true group sizes as N_k = tau_k N^{alpha_k}, with alpha=1 for large groups and alpha<1 for small ones, the K-means least-squares estimators of slopes, grouped fixed effects, and memberships remain consistent and asymptotically normal provided the sample length T grows faster than N^{1-2 alpha_{K0}} (or faster than a power of N when alpha_{K0} >= 1/2) for the smallest group. Smaller groups demand longer panels. For selection, if N^{1-alpha_{m+1}} h_{NT} -> infinity then any information criterion underestimates the number of groups; Bai and Ng's penalty satisfies this when N/T -> i","pith_inferences":["One editorial extension: the rate N^{1-2 alpha} acts as an effective sample size per small-group parameter, so a practitioner could estimate alpha (e.g., from the smallest detected group) and check whether their T exceeds the required growth rate; the theory predicts a sharp threshold where classification suddenly becomes reliable.","A second extension: the under-penalization failure of Bai and Ng's criterion likely transfers to other clustering settings beyond panels—any partition method whose penalty decays like (ln N)/N will drop rare subpopulations when N dominates sample length; the MIC principle (penalty no faster than (ln N)/N when N<=T, and (ln N)/N scaled by 1/T with GFE) may be a portable fix.","A testable extension: run the Monte Carlo with T set exactly at the boundary T = c N^{1-2 alpha} for a small group with alpha = 0.3; the theory predicts that slightly below this boundary the membership error rate fails to vanish, while slightly above it succeeds—this step-like transition can be verified directly."],"forward_implications":["If the central claim is correct, then in panels with small groups, adding more cross-sections N with fixed T does not improve—and can worsen—classification accuracy and parameter estimation for the small groups; only a longer T helps.","Bai and Ng's information criterion can silently miss rare groups whenever N is large relative to T, and in the grouped-fixed-effects model it underestimates the number of groups even when all groups are large.","The BIC-type penalty overestimates the number of groups in finite samples in models without grouped fixed effects, so its use in that setting can create spurious groups.","The proposed MIC penalties (h_MIC1 without GFE, h_MIC2 with GFE) avoid both underestimation and overestimation in simulations, and in the empirical application they detect a 3-firm metal group that Bai and Ng misses while keeping the total group count at a parsimonious three.","The paper's results imply that reports of a group structure should always be accompanied by a check of whether T is large enough relative to the smallest group's size."],"supporting_citations":[{"why":"Supplies the least-squares/K-means estimator, the grouped-fixed-effects model, and the baseline asymptotic results that this paper extends to small groups.","marker":"Bonhomme and Manresa (2015)"},{"why":"Supplies the information criterion and the specific penalty h_BN whose underestimation behavior is analyzed and shown to be inconsistent with small groups.","marker":"Bai and Ng (2002)"},{"why":"Provides the standard consistency argument for number-of-groups estimation via condition (7); this paper shows that condition fails when small groups are present.","marker":"Su, Shi and Phillips (2016)"},{"why":"One of the penalized-objective estimators whose condition for consistently estimating the number of groups does not hold in the presence of small groups.","marker":"Wang, Phillips and Su (2018)"},{"why":"A recent small-group estimator whose consistency condition for the number of groups is shown to fail with small groups, and which does not cover grouped fixed effects.","marker":"Mehrabani (2023)"},{"why":"Motivates the empirical application and supplies the sales-growth model and regressors used in the Japanese firm data illustration.","marker":"Lumsdaine, Okui and Wang (2023)"}],"fun_headline_variants":["Rare panel groups need longer samples for K-means","Modified criterion finds small groups in panel K-means","Small groups in panels: extend time to cluster well","New penalty selects rare subgroups without oversplitting","Panel K-means: small groups dictate sample length"],"cache_read_input_tokens":44800,"weakest_assumption_plain":"The load-bearing premise is Assumption 2(g): the time-series length T must grow faster than N (or faster than N^{1-2 alpha}, where alpha is the exponent of the smallest group's size); if the panel is short relative to the number of units, the proofs of membership consistency and asymptotic normality collapse, and the paper itself flags this as a limitation.","fun_headline_variants_meta":{"raw":{"variants":["Rare panel groups need longer samples for K-means","Modified criterion finds small groups in panel K-means","Small groups in panels: extend time to cluster well","New penalty selects rare subgroups without oversplitting","Panel K-means: small groups dictate sample length"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1236,"prompt_tokens":719,"completion_tokens":517,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":455}},"tokens_in":463,"tokens_out":517,"duration_ms":6617,"temperature":1.0,"reasoning_tokens":455,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:54:24.638433+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a panel with N = 200, T = 10, a small group of size about N^{0.3} (roughly 5 units), well-separated slope parameters, and group sizes otherwise proportional to N. If the proposed MIC selects more than the true K0 groups, or if the K-means membership error rate fails to vanish as N grows with T fixed, the central claim that T must grow like N^{1-2 alpha} would be contradicted. More directly, run the same DGP with T = c N^{0.8} versus T = c N^{1.2} for alpha = 0.3; the theory predicts the first fails to classify the small group perfectly in large samples, while the second succeeds.","supporting_citations":[{"cited_title":"Spectral and post-spectral estimators for grouped panel data models","cited_arxiv_id":"2212.13324","evidence_quote":"Supplies the information criterion and the specific penalty h_BN whose underestimation behavior is analyzed and shown to be inconsistent with small groups."}],"review_version":1}