Pith. sign in

REVIEW 3 major objections 5 minor

Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper claims that a six-week human–AI collaboration tightened both sides of the Grothendieck constant $K_G$, pinning its tenths digit to 7.

desk verdict Solid case study with an honest failure record; the math needs the companion paper to be reviewed along with it. read the letter →

arxiv 2608.11195 v2 pith:XWOFN55P submitted 2026-08-11 cs.AI cs.CCcs.HCmath.FA

classification cs.AIcs.CCcs.HCmath.FA
keywords GrothendieckconstantKrivineschemesSDPintegralitygaproundingalgorithmslowerboundwithouthardinstancelong-horizonAIresearchhuman-AIcollaborationGaussianharmonicanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports that a long-horizon AI research system, steered asynchronously by human operators, tightened the known range of the Grothendieck constant $K_G$ to $\frac{6\pi}{11}\le K_G\le \frac{\pi}{2\log(1+\sqrt{2})}-3.47\times10^{-4}$. The lower bound is the first for this constant that does not exhibit a hard instance; it instead proves that every Krivine rounding scheme obeys the same obstruction. The upper bound comes from a new family of limiting Krivine schemes, the first improvement obtained by letting the dimension of the rounding scheme grow. If the results stand, the constant's tenths digit is 7, and AI systems can contribute mathematically novel steps to long-horizon research while still being unreliable at research judgement and record-keeping.

What carries the argument

The load-bearing object is the normalized correlation function $H(t)=\frac{\pi}{2}\mathbb{E}[f(X)g(Y)]$ of a Krivine scheme, a pair of odd sign functions on $\mathbb{R}^k$ used to round SDP vectors to signs via correlated Gaussians. Its Hermite coefficients $b_1,b_3,\ldots$ control the scheme's rounding quality, and the paper's key identity is the universal affine constraint $b_3\ge 2b_1-\frac{11}{6}$, which is linear in the scheme and therefore passes to mixtures and limits. The Naor-Regev theorem then converts any ceiling of this type into a lower bound on $K_G$. For the upper bound, the cubic-quintic limiting scheme works by making the correlation function closer to a straight line in a higher-dimensional limit, suppressing the nonlinear terms that limit half-space rounding.

What would settle it

Independently recompute, in interval-certified arithmetic, the one-dimensional Gaussian inequalities used in the companion paper's proof that every Krivine scheme satisfies $b_3\ge 2b_1-11/6$; a single certified violation for a mixed or limiting scheme would refute the lower-bound proof.

Watch

Extended reading notes

Core claim

The central mathematical claims are two. First, the correlation function $H(t)=b_1t+b_3t^3+\cdots$ of every Krivine scheme satisfies $b_3\ge 2b_1-\frac{11}{6}$; because this affine inequality survives averaging and limits, the Naor-Regev optimality theorem converts it into $K_G\ge \frac{6\pi}{11}$. Second, an explicit cubic-quintic limiting Krivine scheme, obtained by letting the rounding dimension grow, shows $K_G\le \frac{\pi}{2\log(1+\sqrt{2})}-3.47\times10^{-4}$. The paper presents the discovery narrative: the lower bound was found and first proved by the AI system after human operators redirected it from a plateaued upper-bound search, and was then independently verified by the authors.

Load-bearing premise

The lower bound rests on the assumption that an inequality verified for finitely many rounding schemes survives averaging and limits; the paper records that this transfer was the part repaired during adversarial review, and if the repair is wrong the bound $K_G\ge 6\pi/11$ does not follow.

Editorial extensions

If this is right

  • The interval $\frac{6\pi}{11}\le K_G\le \frac{\pi}{2\log(1+\sqrt{2})}-3.47\times10^{-4}$ determines the tenths digit of $K_G$ as 7, replacing the previous wide interval $1.6769\ldots\le K_G\le 1.7822\ldots$.
  • The lower-bound proof is the first for $K_G$ that does not construct a hard instance; it establishes a universal ceiling on every Krivine scheme, so future lower-bound work in this framework must respect the constraint $b_3\ge 2b_1-\frac{11}{6}$.
  • The cubic-quintic upper bound answers the question of whether higher-dimensional rounding schemes help, giving the first explicit improvement over Krivine's bound obtained by letting the dimension grow.
  • The paper's record indicates that a long-horizon AI research system steered by roughly forty human directives can execute novel proof construction but requires human correction of research judgement and record-keeping, as shown by the session-18 pivot and the session-44 withdrawal.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same template--prove an affine constraint on the leading coefficients of every rounding scheme, then dualize through an optimality theorem--may transfer to other constants whose optimal rounding families are known, yielding lower bounds from failed upper-bound searches.
  • Editorial inference: the paper's diagnosis predicts that training models on process records, such as failed approaches, pivots, and lost caveats, rather than finished proofs should specifically improve their research judgement and memory; this could be tested on long-horizon research logs.
  • Editorial inference: the machine-verified upper bounds $1.781801841033$ and $1.7813319810625639$ were not human-certified; independently checking their certificates would either tighten the interval further or expose a flaw without requiring new mathematics.
  • Editorial inference: one natural extension is to vary the degree of the polynomial boundary beyond cubic-quintic in limiting Krivine schemes and search numerically for larger inverse-majorant parameters, since the paper's machinery appears designed for systematic expansion of the scheme family.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a long-horizon human-AI collaboration aimed at bounding the Grothendieck constant KG. It claims the interval 6π/11 ≤ KG ≤ π/(2 log(1+√2)) − 3.47×10^{-4}, obtained from a new upper-bound construction (limiting Krivine schemes) and a non-constructive lower bound proved via a universal affine constraint on Hermite coefficients, dualized through Naor-Regev's optimality theorem. The mathematical theorems are stated in abridged form with proofs deferred to a companion paper; the main body is a case study of an AI research system, including a 240-session run, cost and timeline data, a methodological framework, and an analysis of the system's strengths and weaknesses. The paper is transparent about the provenance of the results, explicitly reporting that only the lower bound 6π/11 has been human-verified, that several reported bounds are only machine-verified, and that a session-44 upper bound was withdrawn after internal audit.

Significance. If the deferred companion-paper proofs are correct, the reported interval is a substantial tightening of the best known bounds on KG and the lower-bound mechanism is genuinely novel, being the first lower bound not obtained by constructing a hard instance. The case study itself is also valuable: it gives a detailed, quantified account of a long-horizon AI-assisted research program and an honest decomposition of the system's capabilities into technical execution, research judgment, and research-state representation. The paper's transparency about failures—including the withdrawal of a record and the repair of a substantive gap during adversarial review—is a strength. However, the central mathematical claims are not checkable from this submission, and no certificates or code are included, so the significance is currently conditional on the companion paper and on the reliability of an internal verification protocol that the paper itself shows can fail.

major comments (3)
  1. [Section 3, Theorem 3.2] The lower bound KG ≥ 6π/11 depends on Theorem 3.2, but the theorem is stated without proof and deferred to the companion paper [SLX+26]. The proof outline in Section 6 and Appendix A does not give the ternary inequality, the certificate bound, or the weak-* topology used to extend the affine constraint from finite schemes to measurable mixtures and limits. Because Appendix A records that the weak-* passage, a Jensen direction, and the certificate coverage were repaired during adversarial review, the load-bearing step is not checkable from this submission. Please include the full proof, or a complete proof sketch with the certificate, or explicitly reclassify the lower bound as a conditional result proved in [SLX+26] and not established in this paper.
  2. [Section 3, Theorem 3.1] The upper bound is likewise presented without its construction or certificate. The claim that the cubic–quintic scheme yields KG ≤ π/(2 log(1+√2)) − 3.47×10^{-4} is central to the headline interval, but the paper gives only Figure 2 and a one-sentence description. Please either provide the scheme and the verified numerical certificate in an appendix or mark the upper bound as a statement whose proof appears in the companion paper, noting that the construction predates the AI research system studied here.
  3. [Section 5, Table 2] The reported further improvements—27π/49, 51π/92, 1.781801841033, and 1.7813319810625639—are described as machine-verified by the harness's internal protocol, but no certificates, code, or logs are included in the arXiv artifact. Given the session-44 withdrawal, where an uncertified value set a record because a caveat was lost in the research-state representation, the internal verification protocol alone is not a substitute for auditable artifacts. Please either release the certificates and define the protocol operationally, or report these numbers as unverified system outputs rather than as results.
minor comments (5)
  1. [Figure 2] The right-hand panel's boundary formula appears garbled as sgn(z2 (z3 1 3z1)); it should presumably be something like sgn(z2(z1^3 − 3z1)).
  2. [Abstract and Section 3] The abstract reports the upper bound as π/(2 log(1+√2)) − 10^{-4}, while Theorem 3.1 gives the sharper value −3.47×10^{-4}; please use a consistent rounded form or explicitly state the relation between the two.
  3. [References] Reference [KN11] lacks a year, venue, or publisher; please complete the bibliographic entry.
  4. [Section 4] The abstract credits the improvements to 'an AI research system,' but Section 4 states that the cubic–quintic construction predates the system and was found through conversation with GPT-5.5-Pro; please clarify that the upper-bound construction is not part of the harness run.
  5. [Table 2] The follow-up rows are not dated or tied to specific sessions, which makes it difficult to reconstruct the timeline of the withdrawal and the later bounds from the reported archive.

Circularity Check

1 steps flagged · score 2.0 of 10

The lower and upper bounds are deferred to the same-author companion paper [SLX+26]; the math derivation itself is not definitionally circular, so the score stays low.

  1. self citation load bearing [Abstract and Section 3, Theorem 3.2; see also Section 6]
    "The mathematical results are presented and proved in a companion paper [SLX+26]. ... Theorem 3.2 (Lower bound, abridged). The correlation function H(t)=b1t+b3t3+... of every Krivine scheme satisfies the constraint b3 ≥ 2b1 − 11/6. Combined with the optimality theorem of Naor and Regev [NR14], this implies KG ≥ 6π/11 = 1.7135..."

    Within this manuscript, the lower bound KG ≥ 6π/11 is not derived; it is assembled from Theorem 3.2 plus the external Naor–Regev theorem. The proof of Theorem 3.2 — the universal constraint on Krivine schemes — is not present in this arXiv artifact but is cited to [SLX+26], whose author list is a permutation of the present authors. Thus the paper's headline mathematical claim rests on a load-bearing self-citation and cannot be audited from the artifact alone. This is a deferral of proof rather than an equation reducing to itself, and the companion paper is a separate document, so it is a mild, not definitional, circularity.

full rationale

No step in the paper's own derivation chain equates a conclusion to an input by construction. The affine family b3 ≥ (1+λ)b1 − (λ+5/6) is chosen for transportability under mixtures and limits, with λ a proof parameter, not a value fitted to the target bound; at λ=1 it gives Γ1=11/12 and KG ≥ 6π/11 via the external Naor–Regev optimality theorem. The lower-bound mechanism does not rename a known hard-instance construction, and the paper candidly records that the original problem statement already suggested the dualization, so the AI-credit narrative is qualified rather than circular. The main weakness is verifiability: Theorem 3.2 and the cubic–quintic upper bound are imported from [SLX+26], and Appendix A reports that the weak-* mixture passage, a Jensen direction, and the certificate coverage were repaired during adversarial review, none of which can be checked from this artifact. That is a missing-support and verification gap, not a definitional reduction; the score of 2 reflects only the load-bearing same-author companion citation for the headline interval, while the case-study content stands on its own documented run record.

Assumptions & free parameters 1 free parameters · 4 assumptions · 1 invented entities

The central mathematical claim rests on the companion paper's proof, the external Naor-Regev theorem, and the soundness of interval-arithmetic certificates. The only hand-chosen parameter is lambda, set to 1 for provability; it is not fitted to KG. No new physical entities are introduced.

free parameters (1)
  • lambda (affine constraint slope) = 1
    The affine family b3 >= (1+lambda)b1 - (lambda+5/6) is indexed by lambda >= 1/2. The paper chooses lambda = 1 because the proof then reduces to elementary structure, giving Gamma_1 = 11/12 and KG >= 6pi/11. This is a proof choice rather than a fit to KG data.
assumptions (4)
  • standard math Naor-Regev optimality theorem: the optimal approximation ratio of Krivine schemes equals KG.
    Used in Section 6 and Theorem 3.2 to convert a universal ceiling on all schemes into a lower bound on KG; originally proved by Naor and Regev [NR14].
  • domain assumption Limiting Krivine schemes inherit the rounding guarantees of the finite schemes converging to them.
    Defined in Section 2; required so the enlarged search space is valid for rounding. The proof is deferred to the companion paper.
  • ad hoc to paper The affine Hermite constraint extends from finite schemes to measurable mixtures and coefficientwise limits via a weak-* argument.
    Appendix A reports that the weak-* lemma was added after adversarial review; this transfer step is the fragile part that makes the lower bound dualize.
  • domain assumption Interval-arithmetic certificates produced with Arb are sound.
    Section 4 and Appendix A rely on computer-assisted certificates without a formal proof object such as Coq or Lean.
invented entities (1)
  • Limiting Krivine schemes
    purpose: Enlarge the space of rounding algorithms so the cubic-quintic upper bound and the lower-bound obstruction live in a common class.
    A mathematical construction introduced in the companion paper. It is not an empirical entity and has no falsifiable handle outside its proof, but it is a definition rather than an unexplained postulate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration." pith.science (2026). https://pith.science/paper/XWOFN55P

@misc{pith2026260811195,
  author       = {Pith},
  title        = {Pith review of: Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XWOFN55P}},
  note         = {Machine review of arXiv:2608.11195}
}
abstract

AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their continuous relaxations. Specifically, while the precise value of $K_G$ is not known, we recently tightened the best known bounds to \[ \frac{6\pi}{11} \;\le\; K_G \;\le\; \frac{\pi}{2\log(1+\sqrt2)} - 10^{-4}. \] Crucially, these improvements were achieved using an AI research system that could arrive at insights deemed novel by domain experts. We give a detailed discussion of our experience using AI for mathematics research, particularly touching upon its strengths and weaknesses, as well as our experience with creating ideal conditions for AI to arrive at breakthrough insights.

Figures

Figures reproduced from arXiv: 2608.11195 by the authors.

Figure 1
Figure 1. From a hard combinatorial problem to a rounded solution. (a) An instance of the bilinear optimization problem, drawn as a weighted graph: every node carries an unknown sign, every edge a weight aij (an orange +2 edge rewards giving its endpoints equal signs; a dark −2 edge rewards opposite signs), and the goal is to choose the signs maximizing the total reward. (b) The SDP relaxation replaces signs by unit vectors; … view at source ↗
Figure 2
Figure 2. Krivine schemes are partitions of space. A scheme labels the points of R k with ±1 (orange is +1, dark is −1; shown here for k = 2), and each SDP vector is rounded to the label of a correlated random Gaussian point. Left: the half-space partition underlying hyperplane rounding and Krivine’s classical bound. Right: a partition with a cubic boundary, of the kind used by the cubic–quintic scheme of the companion paper;… view at source ↗
Figure 3
Figure 3. A simplified research cycle modelling the process of long-horizon mathematical [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.