REVIEW 3 major objections 5 minor
Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that a six-week human–AI collaboration tightened both sides of the Grothendieck constant $K_G$, pinning its tenths digit to 7.
desk verdict Solid case study with an honest failure record; the math needs the companion paper to be reviewed along with it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the normalized correlation function $H(t)=\frac{\pi}{2}\mathbb{E}[f(X)g(Y)]$ of a Krivine scheme, a pair of odd sign functions on $\mathbb{R}^k$ used to round SDP vectors to signs via correlated Gaussians. Its Hermite coefficients $b_1,b_3,\ldots$ control the scheme's rounding quality, and the paper's key identity is the universal affine constraint $b_3\ge 2b_1-\frac{11}{6}$, which is linear in the scheme and therefore passes to mixtures and limits. The Naor-Regev theorem then converts any ceiling of this type into a lower bound on $K_G$. For the upper bound, the cubic-quintic limiting scheme works by making the correlation function closer to a straight line in a higher-dimensional limit, suppressing the nonlinear terms that limit half-space rounding.
What would settle it
Independently recompute, in interval-certified arithmetic, the one-dimensional Gaussian inequalities used in the companion paper's proof that every Krivine scheme satisfies $b_3\ge 2b_1-11/6$; a single certified violation for a mixed or limiting scheme would refute the lower-bound proof.
Extended reading notes
Core claim
The central mathematical claims are two. First, the correlation function $H(t)=b_1t+b_3t^3+\cdots$ of every Krivine scheme satisfies $b_3\ge 2b_1-\frac{11}{6}$; because this affine inequality survives averaging and limits, the Naor-Regev optimality theorem converts it into $K_G\ge \frac{6\pi}{11}$. Second, an explicit cubic-quintic limiting Krivine scheme, obtained by letting the rounding dimension grow, shows $K_G\le \frac{\pi}{2\log(1+\sqrt{2})}-3.47\times10^{-4}$. The paper presents the discovery narrative: the lower bound was found and first proved by the AI system after human operators redirected it from a plateaued upper-bound search, and was then independently verified by the authors.
Load-bearing premise
The lower bound rests on the assumption that an inequality verified for finitely many rounding schemes survives averaging and limits; the paper records that this transfer was the part repaired during adversarial review, and if the repair is wrong the bound $K_G\ge 6\pi/11$ does not follow.
Editorial extensions
If this is right
- The interval $\frac{6\pi}{11}\le K_G\le \frac{\pi}{2\log(1+\sqrt{2})}-3.47\times10^{-4}$ determines the tenths digit of $K_G$ as 7, replacing the previous wide interval $1.6769\ldots\le K_G\le 1.7822\ldots$.
- The lower-bound proof is the first for $K_G$ that does not construct a hard instance; it establishes a universal ceiling on every Krivine scheme, so future lower-bound work in this framework must respect the constraint $b_3\ge 2b_1-\frac{11}{6}$.
- The cubic-quintic upper bound answers the question of whether higher-dimensional rounding schemes help, giving the first explicit improvement over Krivine's bound obtained by letting the dimension grow.
- The paper's record indicates that a long-horizon AI research system steered by roughly forty human directives can execute novel proof construction but requires human correction of research judgement and record-keeping, as shown by the session-18 pivot and the session-44 withdrawal.
Reading between the lines
- Editorial inference: the same template--prove an affine constraint on the leading coefficients of every rounding scheme, then dualize through an optimality theorem--may transfer to other constants whose optimal rounding families are known, yielding lower bounds from failed upper-bound searches.
- Editorial inference: the paper's diagnosis predicts that training models on process records, such as failed approaches, pivots, and lost caveats, rather than finished proofs should specifically improve their research judgement and memory; this could be tested on long-horizon research logs.
- Editorial inference: the machine-verified upper bounds $1.781801841033$ and $1.7813319810625639$ were not human-certified; independently checking their certificates would either tighten the interval further or expose a flaw without requiring new mathematics.
- Editorial inference: one natural extension is to vary the degree of the polynomial boundary beyond cubic-quintic in limiting Krivine schemes and search numerically for larger inverse-majorant parameters, since the paper's machinery appears designed for systematic expansion of the scheme family.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a long-horizon human-AI collaboration aimed at bounding the Grothendieck constant KG. It claims the interval 6π/11 ≤ KG ≤ π/(2 log(1+√2)) − 3.47×10^{-4}, obtained from a new upper-bound construction (limiting Krivine schemes) and a non-constructive lower bound proved via a universal affine constraint on Hermite coefficients, dualized through Naor-Regev's optimality theorem. The mathematical theorems are stated in abridged form with proofs deferred to a companion paper; the main body is a case study of an AI research system, including a 240-session run, cost and timeline data, a methodological framework, and an analysis of the system's strengths and weaknesses. The paper is transparent about the provenance of the results, explicitly reporting that only the lower bound 6π/11 has been human-verified, that several reported bounds are only machine-verified, and that a session-44 upper bound was withdrawn after internal audit.
Significance. If the deferred companion-paper proofs are correct, the reported interval is a substantial tightening of the best known bounds on KG and the lower-bound mechanism is genuinely novel, being the first lower bound not obtained by constructing a hard instance. The case study itself is also valuable: it gives a detailed, quantified account of a long-horizon AI-assisted research program and an honest decomposition of the system's capabilities into technical execution, research judgment, and research-state representation. The paper's transparency about failures—including the withdrawal of a record and the repair of a substantive gap during adversarial review—is a strength. However, the central mathematical claims are not checkable from this submission, and no certificates or code are included, so the significance is currently conditional on the companion paper and on the reliability of an internal verification protocol that the paper itself shows can fail.
major comments (3)
- [Section 3, Theorem 3.2] The lower bound KG ≥ 6π/11 depends on Theorem 3.2, but the theorem is stated without proof and deferred to the companion paper [SLX+26]. The proof outline in Section 6 and Appendix A does not give the ternary inequality, the certificate bound, or the weak-* topology used to extend the affine constraint from finite schemes to measurable mixtures and limits. Because Appendix A records that the weak-* passage, a Jensen direction, and the certificate coverage were repaired during adversarial review, the load-bearing step is not checkable from this submission. Please include the full proof, or a complete proof sketch with the certificate, or explicitly reclassify the lower bound as a conditional result proved in [SLX+26] and not established in this paper.
- [Section 3, Theorem 3.1] The upper bound is likewise presented without its construction or certificate. The claim that the cubic–quintic scheme yields KG ≤ π/(2 log(1+√2)) − 3.47×10^{-4} is central to the headline interval, but the paper gives only Figure 2 and a one-sentence description. Please either provide the scheme and the verified numerical certificate in an appendix or mark the upper bound as a statement whose proof appears in the companion paper, noting that the construction predates the AI research system studied here.
- [Section 5, Table 2] The reported further improvements—27π/49, 51π/92, 1.781801841033, and 1.7813319810625639—are described as machine-verified by the harness's internal protocol, but no certificates, code, or logs are included in the arXiv artifact. Given the session-44 withdrawal, where an uncertified value set a record because a caveat was lost in the research-state representation, the internal verification protocol alone is not a substitute for auditable artifacts. Please either release the certificates and define the protocol operationally, or report these numbers as unverified system outputs rather than as results.
minor comments (5)
- [Figure 2] The right-hand panel's boundary formula appears garbled as sgn(z2 (z3 1 3z1)); it should presumably be something like sgn(z2(z1^3 − 3z1)).
- [Abstract and Section 3] The abstract reports the upper bound as π/(2 log(1+√2)) − 10^{-4}, while Theorem 3.1 gives the sharper value −3.47×10^{-4}; please use a consistent rounded form or explicitly state the relation between the two.
- [References] Reference [KN11] lacks a year, venue, or publisher; please complete the bibliographic entry.
- [Section 4] The abstract credits the improvements to 'an AI research system,' but Section 4 states that the cubic–quintic construction predates the system and was found through conversation with GPT-5.5-Pro; please clarify that the upper-bound construction is not part of the harness run.
- [Table 2] The follow-up rows are not dated or tied to specific sessions, which makes it difficult to reconstruct the timeline of the withdrawal and the later bounds from the reported archive.
Circularity Check
The lower and upper bounds are deferred to the same-author companion paper [SLX+26]; the math derivation itself is not definitionally circular, so the score stays low.
-
self citation load bearing
[Abstract and Section 3, Theorem 3.2; see also Section 6]
"The mathematical results are presented and proved in a companion paper [SLX+26]. ... Theorem 3.2 (Lower bound, abridged). The correlation function H(t)=b1t+b3t3+... of every Krivine scheme satisfies the constraint b3 ≥ 2b1 − 11/6. Combined with the optimality theorem of Naor and Regev [NR14], this implies KG ≥ 6π/11 = 1.7135..."
Within this manuscript, the lower bound KG ≥ 6π/11 is not derived; it is assembled from Theorem 3.2 plus the external Naor–Regev theorem. The proof of Theorem 3.2 — the universal constraint on Krivine schemes — is not present in this arXiv artifact but is cited to [SLX+26], whose author list is a permutation of the present authors. Thus the paper's headline mathematical claim rests on a load-bearing self-citation and cannot be audited from the artifact alone. This is a deferral of proof rather than an equation reducing to itself, and the companion paper is a separate document, so it is a mild, not definitional, circularity.
full rationale
No step in the paper's own derivation chain equates a conclusion to an input by construction. The affine family b3 ≥ (1+λ)b1 − (λ+5/6) is chosen for transportability under mixtures and limits, with λ a proof parameter, not a value fitted to the target bound; at λ=1 it gives Γ1=11/12 and KG ≥ 6π/11 via the external Naor–Regev optimality theorem. The lower-bound mechanism does not rename a known hard-instance construction, and the paper candidly records that the original problem statement already suggested the dualization, so the AI-credit narrative is qualified rather than circular. The main weakness is verifiability: Theorem 3.2 and the cubic–quintic upper bound are imported from [SLX+26], and Appendix A reports that the weak-* mixture passage, a Jensen direction, and the certificate coverage were repaired during adversarial review, none of which can be checked from this artifact. That is a missing-support and verification gap, not a definitional reduction; the score of 2 reflects only the load-bearing same-author companion citation for the headline interval, while the case-study content stands on its own documented run record.
Assumptions & free parameters
free parameters (1)
- lambda (affine constraint slope) =
1
assumptions (4)
- standard math Naor-Regev optimality theorem: the optimal approximation ratio of Krivine schemes equals KG.
- domain assumption Limiting Krivine schemes inherit the rounding guarantees of the finite schemes converging to them.
- ad hoc to paper The affine Hermite constraint extends from finite schemes to measurable mixtures and coefficientwise limits via a weak-* argument.
- domain assumption Interval-arithmetic certificates produced with Arb are sound.
invented entities (1)
-
Limiting Krivine schemes
Cite this review
Pith. "Pith review of Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration." pith.science (2026). https://pith.science/paper/XWOFN55P
@misc{pith2026260811195,
author = {Pith},
title = {Pith review of: Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration},
year = {2026},
howpublished = {\url{https://pith.science/paper/XWOFN55P}},
note = {Machine review of arXiv:2608.11195}
}
abstract
AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations. Specifically, while the precise value of $K_G$ is not known, we recently tightened the best known bounds to \[ \frac{6\pi}{11} \;\le\; K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}. \] Crucially, these improvements were achieved using an AI research system that could arrive at insights deemed novel by domain experts. We give a detailed discussion of our experience using AI for mathematics research, particularly touching upon its strengths and weaknesses, as well as our experience with creating ideal conditions for AI to arrive at breakthrough insights.
Figures
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.