{"id":"1d10ff6b-320a-4347-aa7b-9668a3ed1525","arxiv_id":"2607.17090","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"For shallow polynomial networks over finite fields, the representable-function count is computed exactly for (n,1,k) and (n,2,k), and collapses to a linear-matrix count when the activation degree is a power of the field characteristic.","lead":"Neural networks with polynomial activations become finite sets of functions when weights are chosen from a finite field. This paper counts those sets for several simple architectures, and finds a case where finite fields give only about half the ambient space while complex networks fill it densely.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the m=2 counting is sound; only minor presentation/table errors found.","rationale":"The reader's weakest-assumption identification is correct and the proof of that assumption is sound: for p∤r and r>1, any non-trivial linear combination of two distinct rank-1 symmetric tensors retains a non-zero mixed monomial coefficient, so it cannot collapse to rank 1. This validates the disjoint partition in Proposition 4.3. I checked the (2,2,2) numerical table against the proposition for q=3,5,7 and found exact agreement; the apparent anomaly at p=11 is a relabelling of the q=13 row, since the listed ambient size is 13^6. Lemma 4.2 also appears to have a missing bar on \\bar{M} and/or a missing +1 in the display, but the surrounding text and Table 1 give the correct projective-to-affine conversion. Neither issue undermines the central claim that expressivity is exactly quantified by neuromanifold cardinality for the named architectures. The paper's core contribution is sound and the ACCEPT verdict remains appropriate.","tokens_in":20163,"tokens_out":30523,"duration_ms":272257,"concrete_test":"Recompute Table 2 row labelled p=11 using Proposition 4.3: for q=11 the formula gives |M_{(2,2,2),2}|=887,161 and ambient size 11^6=1,771,561, whereas the values shown (2,415,673 and 4,826,809) are exactly the q=13 outputs. A direct enumeration for n=k=2, r=2, q=11 — or a symbolic count over F_11 — settles the mislabelling and confirms the formula.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central counting claims hold under scrutiny. The critical uniqueness step in Prop. 4.3 is valid: for p∤r and r>1, a non-trivial combination c1*L1^{⊗r}+c2*L2^{⊗r} has a non-zero mixed component (e.g., r*v1^{r-1}*v2 in a basis where L1=e1, L2=e2), so it cannot equal a rank-1 tensor L3^{⊗r}; hence each 2-dimensional span has a unique generating pair and the disjoint partition is justified. I re-derived the (2,2,2) counts from Prop. 4.3 and reproduced the table entries for q=3,5,7. Two minor presentation issues do not affect the theorems: Lemma 4.2 states a projective cardinality but writes |M| instead of |\\bar{M}| (the +1 correction of Lemma 4.1 is missing in the display), and Table 2's row labelled p=11 actually contains the q=13 values (ambient 4,826,809=13^6 and the listed |M| matches Prop. 4.3 at q=13). These are typographical, not mathematical.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an algebraic notion of expressivity for shallow polynomial neural networks over finite fields, measured by the cardinality of the neuromanifold (the image of the parameter map in a product of polynomial rings). It proves a general upper bound via the multinomial congruence in characteristic p, computes exact counts for single-output square activation (d=(n,m,1), r=2) using MacWilliams' enumeration of symmetric matrices, and obtains projective counting formulas for d=(n,1,k) and d=(n,2,k) for arbitrary monomial degrees. The key step for m=2 is a disjoint-partition argument based on the uniqueness of generating pairs of rank-1 tensors when r>1 and p∤r. For degrees divisible by p, a Frobenius reduction equates the count with the p-free part. The paper also highlights a contrast with the complex case: for d=(2,2,2), r=2, the finite-field neuromanifold occupies roughly half of the ambient space, whereas over C it is Zariski dense.","tokens_in":20402,"tokens_out":23258,"duration_ms":179210,"significance":"The results provide the first systematic point counts for neuromanifolds of shallow polynomial neural networks over finite fields, linking network expressivity to classical finite-field enumeration. The counting arguments are rigorous; the critical uniqueness step in Proposition 4.3 is valid, and the Frobenius reduction for p|r is correct. The paper is transparent about limitations: Conjecture 3.6 is stated as a conjecture supported only by four data points, and Remark 4.5 explicitly notes that the m=2 partition does not extend to m≥3. The contrast between finite-field and complex behavior is interesting and may motivate further arithmetic study of neural network expressivity.","major_comments":[],"minor_comments":[{"comment":"The statement writes |\\mathfrak{M}_{d,r}|, but the proof establishes the projective cardinality |\\bar{\\mathfrak{M}}_{d,r}| = |\\mathbb{P}^{n-1}| |\\mathbb{P}^{k-1}|. The affine cardinality is (q-1) times this plus 1 by Lemma 4.1. The same projective/affine confusion appears in Table 1, row (n,1,k). Please correct the statement and table.","section":"Section 4, Lemma 4.2"},{"comment":"The row labelled p=11 has ambient size 4,826,809 = 13^6, and the listed |M| matches Proposition 4.3 at q=13, not q=11. The label should be q=13 (or the row recomputed).","section":"Table 2"},{"comment":"V2 is defined as the intersection over all (i,j,k) of the vanishing of the (j,k)-th minor determinants of A_i, but the set {det A_1 = det A_2 = 0} is not equal to that intersection. Also, V1 uses A_1^{-1}A_2, which is not a polynomial condition when A_1 is singular. The proof should be repaired by defining the bad locus as {det A_1 det A_2 = 0} ∪ {discriminant of charpoly(A_1^{-1}A_2) = 0} after clearing denominators. This is a local issue and does not affect the finite-field counting results.","section":"Section 3.2, proof of Theorem 3.5"},{"comment":"The dichotomy 'r ≠ p^i' vs 'r = p^i' is not correct when r has a p-free factor i>1 and p|r (e.g., r=2p with p odd). The intended reduction via Corollary 5.1.1 is: if the p-free part of r is 1, use Equation (14); otherwise use Equation (13) with the p-free part.","section":"Corollary 4.5.1"},{"comment":"There are minor typesetting issues, e.g., the running header 'M. Zubkov et al.:Preprint submitted to ElsevierPage 1 of 16' and inconsistent use of double/triple bars for projective cardinalities. Please harmonize notation. Also, Corollary 5.3.1's 'min' is redundant when m≥min(n,k); consider simplifying.","section":"General presentation"}],"recommendation":"minor_revision","confidential_remarks":"The central counting results are sound; the issues are local. The proof of Theorem 3.5, while not affecting the main finite-field results, should be corrected before publication. The paper is honest about the unproven conjecture and the m≥3 limitation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you work on quantized networks or finite-field analogues of algebraic geometry in ML. The paper gives the first exact counts of neuromanifolds for shallow polynomial networks over finite fields, for m=1 and m=2 hidden units, plus a reduction theorem when the activation degree is divisible by the field characteristic. I checked the key step: for r>1 and p∤r, a non-trivial combination of two rank-1 symmetric tensors is not rank-1, so the disjoint partition in Prop. 4.3 is valid. I also reproduced the numerical table for (2,2,2) at q=3,5,7 (and, it turns out, 13) from the formulas.\n\nWhat is best: the paper is honest. It labels Conjecture 3.6 as a conjecture, states the range of applicability clearly, and gives a clean proof of the p|r reduction using perfectness of finite fields. The counting arguments are elementary but careful, and the main theorem for m=2 is a real result, not a restatement of known facts.\n\nSoft spots, in order of severity. First, the abstract and highlights lean on the Weil conjectures, but the paper never uses them—the proofs are elementary combinatorics over finite fields. That framing should be dialed down. Second, the 'approximately half' claim for the (2,2,2) architecture is only a conjecture with numerical evidence, not a theorem; the abstract currently makes it sound proven. Third, the highlights claim a 'general lower bound', but what the text actually proves is a lower bound via SDC tuples, not a general formula. That should be reworded.\n\nThere are also two typos that would matter to a reader: Lemma 4.2 states an affine cardinality but proves a projective one (the statement is missing the +1 from Lemma 4.1), and Table 2 labels the last row p=11 but the numbers are for q=13. Both are easy fixes and do not affect any theorem.\n\nWho this is for: people in neuroalgebraic geometry, tensor expressivity, or finite-field counting problems. The results are modest in scope but new, and the proof strategy is transferable. I would send it to a serious referee; it deserves revision, not rejection.","headline":"A solid, honest counting paper that opens a genuinely new finite-field direction in neuroalgebraic geometry; the main formulas are correct and the flaws are typographical or presentational.","tokens_in":20932,"tokens_out":3367,"would_cite":true,"duration_ms":33353,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["14G15","14G05","05A30"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that over a finite field, the expressivity of a shallow polynomial neural network is exactly the number of distinct output tuples its weights can produce—and for several architectures that number is given by explicit formul","keywords":["polynomial neural networks","finite fields","neuromanifold","expressivity","symmetric tensors","matrix rank over finite fields","point counting","characteristic p"],"falsifier":"For a small case such as d=(2,2,2), r=2 over F_5, enumerate all 5^8 weight assignments, compute the resulting pair of quadratic forms, and count distinct pairs; the count should equal 7945 as reported in the paper's Table 2. A mismatch would refute Proposition 4.3's partition argument. As a sharper test, compute the exact cardinality over F_13 and check whether it is near half of 13^6, which the conjectured limit would predict.","tokens_in":20019,"feed_emoji":"🧮","tokens_out":6789,"duration_ms":58095,"temperature":0.7,"pith_summary":"This paper establishes that the expressivity of a shallow polynomial neural network over a finite field is precisely the number of distinct output tuples its parameter map can produce—the cardinality of its neuromanifold—and it computes this number exactly for several architectures. For a single hidden unit, for two hidden units, and for activation degrees divisible by the field's characteristic, the counts are given by closed formulas built from matrices of bounded rank and from projective-line counts. The authors also prove a general upper bound in terms of the base-p expansion of the activation degree. The most striking consequence is a demonstrated split: over the complex numbers a certain shallow quadratic network generates a Zariski-dense set of functions, while over finite fields it occupies only about half of the ambient space. If the paper is right, expressivity is not a purely architectural property; the characteristic of the coefficient field can halve it.","feed_headline":"Over finite fields, a small net covers only half the function space","feed_subtitle":"Exact point counts show a quadratic network's expressivity drops to ~50 percent when coefficients live in F_p, not C.","key_machinery":"The central object is the parameter map sending a weight matrix pair to the tuple of homogeneous polynomials (sum_s b_{is}(a_{s1}x_1+...+a_{sn}x_n)^r)_{i=1..k}; its image is the neuromanifold. The counting machinery combines four ingredients: a multinomial congruence theorem showing which monomial coefficients can be nonzero modulo p, giving the upper bound q^{γ_{n,p}(r)k}; the classical formula for the number of symmetric matrices of given rank over a finite field, used for the square-activation single-output case; projective-space counting, which turns the neuromanifold into a product P^{n-1} × P^{k-1} for m=1 and into a union of fibers over rank-1 tensor pairs for m=2; and the perfection","core_discovery":"The central claim is that, for architectures (n,1,k), (n,2,k), and for activation degrees divisible by p, the neuromanifold's cardinality is given by exact formulas. In particular, for d=(2,2,2) with r=2, the finite-field neuromanifold has cardinality asymptotically half of the ambient space as p grows, whereas the same architecture over the complex numbers has a Zariski-dense image. The proof rests on a partition of the neuromanifold by the dimension of the span of the output tuple and on the fact that, when r>1, a 2-dimensional subspace of symmetric tensors contains at most two projectivized rank-1 tensors, so each representable tuple has a unique generating pair of rank-1 forms.","pith_inferences":["The 1/2 limit for (2,2,2), r=2 suggests a pattern: events that are 'probability zero' over the complex numbers, such as repeated eigenvalues of a pencil of symmetric matrices, become events of positive density over finite fields; quantifying this density for other architectures could predict when low-precision networks lose expressivity.","The exact enumeration makes a concrete test possible for quantized networks: for small q one can brute-force all weight assignments and compare the representable function count against the formulas, giving a training-free measure of capacity.","The equality for p-multiple degrees implies that in characteristic p, monomial activations of degree r and r/p behave identically; this could guide hardware designers choosing between activation degrees in low-precision settings.","A natural extension left open by the paper is the m≥3 case, where the unique-generating-pair argument fails; incidence-counting of decompositions into three or more rank-1 tensors would be needed to extend the exact formulas."],"forward_implications":["If the formulas hold, expressivity over a finite field becomes an exact, computable number: for a given architecture one can enumerate every representable k-tuple of homogeneous polynomials without training.","Over the complex numbers, the (n,n,2), r=2 network fills the ambient space in the Zariski sense, but over F_p its arithmetic capacity tends to 1/2 for n=2; the same architecture is thus expressive or half-expressive depending on the coefficient field.","Activation degrees that differ by a factor of p are equally expressive over F_q: the neuromanifold cardinality is unchanged when r is replaced by r/p, because every field element is a p-th power.","When the activation degree is a power of p, the neuromanifold coincides exactly with the set of k×n matrices of rank at most m, reducing the network to a linear map with a rank constraint.","The general upper bound q^{γ_{n,p}(r)k} means that the base-p digits of r, not just its size, control what a shallow network can represent."],"fun_headline_variants":["Finite fields halve shallow net expressivity","Over F_p, small nets cover only half the function space","Exact counts: finite field nets hit half the ambient space","Shallow net over finite field: half the functions, precisely","Field characteristic cuts shallow net expressivity to 50%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The counting formulas for two hidden units assume that, when the activation degree exceeds 1, a two-dimensional subspace of symmetric tensors contains at most two projectivized rank-1 tensors, so that every representable tuple has a unique generating pair; if a subspace ever contained three such rank-1 directions, the partition would overcount and the closed-form counts would be wrong.","fun_headline_variants_meta":{"raw":{"variants":["Finite fields halve shallow net expressivity","Over F_p, small nets cover only half the function space","Exact counts: finite field nets hit half the ambient space","Shallow net over finite field: half the functions, precisely","Field characteristic cuts shallow net expressivity to 50%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1351,"prompt_tokens":669,"completion_tokens":682,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":413,"completion_tokens_details":{"reasoning_tokens":600}},"tokens_in":413,"tokens_out":682,"duration_ms":5747,"temperature":1.0,"reasoning_tokens":600,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T19:04:28.870473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a small case such as d=(2,2,2), r=2 over F_5, enumerate all 5^8 weight assignments, compute the resulting pair of quadratic forms, and count distinct pairs; the count should equal 7945 as reported in the paper's Table 2. A mismatch would refute Proposition 4.3's partition argument. As a sharper test, compute the exact cardinality over F_13 and check whether it is near half of 13^6, which the conjectured limit would predict.","supporting_citations":[],"review_version":1}