{"id":"b68764db-7ca3-4438-ae4a-85cbc57d2f89","arxiv_id":"2608.11195","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A case study showing an AI research system, with human steering, helped prove new bounds 6pi/11 <= KG <= pi/(2 log(1+sqrt(2))) - 3.47e-4 on the Grothendieck constant.","lead":"This paper documents a six-week human-AI research program that tightened known bounds on the Grothendieck constant, a long-standing open problem in mathematics. It analyzes where the AI system excelled and where human judgment was still required, providing a template for long-horizon AI-assisted research.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lower bound 6π/11 hinges on the companion-paper proof of the universal constraint b3 ≥ 2b1 − 11/6, which is absent here and was repaired twice during adversarial review; a residual gap in the weak-* limit or the ternary inequality would invalidate the main theorem.","rationale":"The Reader's weakest_assumption is exactly the correctness of the companion-paper proof of the universal coefficient inequality and its extension to mixtures and limits. I agree that this is the most load-bearing premise: the lower bound 6π/11 is new, non-constructive, and determines the tenths digit of KG, and it rests entirely on Theorem 3.2 plus the Naor–Regev dualization. The paper is honest about the fragility of the proof: Appendix A explicitly says the chain had a substantive weak-* gap, a Jensen-direction error, and a certificate coverage defect, all repaired during adversarial review. That is a red flag that the remaining proof is delicate, and the abridged presentation plus deferred companion paper makes independent assessment impossible from this artifact. My concern is not that the theorem is false; it is that the presented submission does not contain enough evidence to move from 'conditional' to 'accepted,' and no amount of AI-methodology analysis can substitute for a certified companion proof. The suggested test targets the exact steps that were repaired: the ternary inequality reduction, the certificate margin, and the weak-* continuity. Because the Reader already assigned CONDITIONAL and my analysis does not identify a demonstrated flaw, I recommend the verdict remain unchanged.","tokens_in":12368,"tokens_out":24935,"duration_ms":210485,"concrete_test":"Obtain [SLX+26] and independently re-derive the λ=1 case: write h=(f+g)/2, k=(f−g)/2, and verify that the claimed equivalence between b3 ≥ 2b1 − 11/6 and the stated one-dimensional ternary inequality is correct for all odd ±1 functions f,g; then reproduce the interval certificate with independent Arb code for the finite set of polynomial inequalities; finally, write out the weak-* limit lemma explicitly and confirm that b1,b3 are continuous linear functionals on the closed convex hull of finite schemes, so the affine constraint passes to measurable mixtures and limits. If any of these steps fails on an explicit scheme, Theorem 3.2 and the lower bound 6π/11 do not follow.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central new mathematical result is Theorem 3.2: every Krivine scheme (in the limiting sense needed for Naor–Regev) satisfies b3 ≥ 2b1 − 11/6. The entire lower bound KG ≥ 6π/11 is the dual of this statement via [NR14]. The proof is not in this submission; it is deferred to [SLX+26], and the paper's own Appendix A records that the chain 'found and repaired the one substantive gap (the passage from finitely many schemes to measurable mixtures, closed by a weak-* lemma), a Jensen inequality applied in the wrong direction, and a small defect in the interval certificate's coverage.' These are exactly the steps that must hold for the universal quantification. The affine inequality is preserved under convex mixtures and coefficientwise limits only if (i) the Hermite coefficient functionals b1,b3 are continuous on the closure of the finite scheme class in the topology used for limiting schemes, and (ii) the reduction to the one-dimensional 'ternary inequality' (session 57) has no hidden regularity assumption on f,g. The paper does not state the ternary inequality, the certificate bound, or the weak-* topology, and no code or certificate files are included, so none of these can be checked from the arXiv artifact. Since this is the first non-constructive lower bound on KG, any gap here invalidates the paper's headline interval, whatever the merits of the AI case study.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a long-horizon human-AI collaboration aimed at bounding the Grothendieck constant KG. It claims the interval 6π/11 ≤ KG ≤ π/(2 log(1+√2)) − 3.47×10^{-4}, obtained from a new upper-bound construction (limiting Krivine schemes) and a non-constructive lower bound proved via a universal affine constraint on Hermite coefficients, dualized through Naor-Regev's optimality theorem. The mathematical theorems are stated in abridged form with proofs deferred to a companion paper; the main body is a case study of an AI research system, including a 240-session run, cost and timeline data, a methodological framework, and an analysis of the system's strengths and weaknesses. The paper is transparent about the provenance of the results, explicitly reporting that only the lower bound 6π/11 has been human-verified, that several reported bounds are only machine-verified, and that a session-44 upper bound was withdrawn after internal audit.","tokens_in":12665,"tokens_out":6564,"duration_ms":58003,"significance":"If the deferred companion-paper proofs are correct, the reported interval is a substantial tightening of the best known bounds on KG and the lower-bound mechanism is genuinely novel, being the first lower bound not obtained by constructing a hard instance. The case study itself is also valuable: it gives a detailed, quantified account of a long-horizon AI-assisted research program and an honest decomposition of the system's capabilities into technical execution, research judgment, and research-state representation. The paper's transparency about failures—including the withdrawal of a record and the repair of a substantive gap during adversarial review—is a strength. However, the central mathematical claims are not checkable from this submission, and no certificates or code are included, so the significance is currently conditional on the companion paper and on the reliability of an internal verification protocol that the paper itself shows can fail.","major_comments":[{"comment":"The lower bound KG ≥ 6π/11 depends on Theorem 3.2, but the theorem is stated without proof and deferred to the companion paper [SLX+26]. The proof outline in Section 6 and Appendix A does not give the ternary inequality, the certificate bound, or the weak-* topology used to extend the affine constraint from finite schemes to measurable mixtures and limits. Because Appendix A records that the weak-* passage, a Jensen direction, and the certificate coverage were repaired during adversarial review, the load-bearing step is not checkable from this submission. Please include the full proof, or a complete proof sketch with the certificate, or explicitly reclassify the lower bound as a conditional result proved in [SLX+26] and not established in this paper.","section":"Section 3, Theorem 3.2"},{"comment":"The upper bound is likewise presented without its construction or certificate. The claim that the cubic–quintic scheme yields KG ≤ π/(2 log(1+√2)) − 3.47×10^{-4} is central to the headline interval, but the paper gives only Figure 2 and a one-sentence description. Please either provide the scheme and the verified numerical certificate in an appendix or mark the upper bound as a statement whose proof appears in the companion paper, noting that the construction predates the AI research system studied here.","section":"Section 3, Theorem 3.1"},{"comment":"The reported further improvements—27π/49, 51π/92, 1.781801841033, and 1.7813319810625639—are described as machine-verified by the harness's internal protocol, but no certificates, code, or logs are included in the arXiv artifact. Given the session-44 withdrawal, where an uncertified value set a record because a caveat was lost in the research-state representation, the internal verification protocol alone is not a substitute for auditable artifacts. Please either release the certificates and define the protocol operationally, or report these numbers as unverified system outputs rather than as results.","section":"Section 5, Table 2"}],"minor_comments":[{"comment":"The right-hand panel's boundary formula appears garbled as sgn(z2 (z3 1 3z1)); it should presumably be something like sgn(z2(z1^3 − 3z1)).","section":"Figure 2"},{"comment":"The abstract reports the upper bound as π/(2 log(1+√2)) − 10^{-4}, while Theorem 3.1 gives the sharper value −3.47×10^{-4}; please use a consistent rounded form or explicitly state the relation between the two.","section":"Abstract and Section 3"},{"comment":"Reference [KN11] lacks a year, venue, or publisher; please complete the bibliographic entry.","section":"References"},{"comment":"The abstract credits the improvements to 'an AI research system,' but Section 4 states that the cubic–quintic construction predates the system and was found through conversation with GPT-5.5-Pro; please clarify that the upper-bound construction is not part of the harness run.","section":"Section 4"},{"comment":"The follow-up rows are not dated or tied to specific sessions, which makes it difficult to reconstruct the timeline of the withdrawal and the later bounds from the reported archive.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The case-study contribution is informative and unusually candid, but the mathematical headline depends on a companion paper that is not part of this submission and on internally certified computations that are not auditable from the arXiv artifact. If the editor accepts the mathematical claims as external to this paper, the case-study value is sufficient for publication after the clarity revisions; otherwise the manuscript needs to include the proof and certificate material, or the claims must be presented as conditional on the companion paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this paper as a case study, not as the proof. The main mathematical theorems are abridged, and the arguments live in a companion paper that this submission does not include. The significant thing here is the documentation: a roughly 240-session human-AI research run with telemetry, a withdrawn upper bound, and a lower-bound proof that survived adversarial review only after fixing a weak-* gap, a reversed Jensen inequality, and a certificate-coverage defect. That honesty is real and rare.\n\nThe new math is the first lower bound on KG that does not construct a hard instance, using the Naor-Regev dualization and an affine constraint on Hermite coefficients. If the companion paper is right, this is important. But I cannot verify it from this text. The stress-test concerns are legitimate: the weak-* limiting argument and the ternary inequality are load-bearing and they are asserted, not shown. The paper itself flags these as the repaired gaps, so the authors know where the pressure points are.\n\nThe circularity worry is mostly unfounded. The parameter lambda parametrizes a family of proof obligations; it is not fitted to the target. The derivation is not circular in the sense that matters.\n\nAlso note: the authors concede that the central dualization idea was suggested in the problem statement, and the run's critical pivot was a human directive. So the AI's claimed 'insight' is more accurately described as strong execution on a human-provided framework. That moderates the novelty but does not destroy it—the system produced a complete proof and identified the right sub-lemmas.\n\nFor a careful reader this is a valuable contribution to the methodology of AI-assisted mathematics. It deserves peer review, but the companion paper must be part of the reviewed artifact, and the authors should be asked to include the weak-* lemma and ternary inequality or provide a downloadable certificate.\n\nIf I had to decide: send it to review, with the companion. Don't let the case study stand alone.","headline":"Solid case study with an honest failure record; the math needs the companion paper to be reviewed along with it.","tokens_in":13205,"tokens_out":2797,"would_cite":true,"duration_ms":25486,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a six-week human–AI collaboration tightened both sides of the Grothendieck constant $K_G$, pinning its tenths digit to 7.","keywords":["Grothendieck constant","Krivine schemes","SDP integrality gap","rounding algorithms","lower bound without hard instance","long-horizon AI research","human-AI collaboration","Gaussian harmonic analysis"],"falsifier":"Independently recompute, in interval-certified arithmetic, the one-dimensional Gaussian inequalities used in the companion paper's proof that every Krivine scheme satisfies $b_3\\ge 2b_1-11/6$; a single certified violation for a mixed or limiting scheme would refute the lower-bound proof.","tokens_in":12156,"feed_emoji":"🤖","tokens_out":9988,"duration_ms":84959,"temperature":0.7,"pith_summary":"The paper reports that a long-horizon AI research system, steered asynchronously by human operators, tightened the known range of the Grothendieck constant $K_G$ to $\\frac{6\\pi}{11}\\le K_G\\le \\frac{\\pi}{2\\log(1+\\sqrt{2})}-3.47\\times10^{-4}$. The lower bound is the first for this constant that does not exhibit a hard instance; it instead proves that every Krivine rounding scheme obeys the same obstruction. The upper bound comes from a new family of limiting Krivine schemes, the first improvement obtained by letting the dimension of the rounding scheme grow. If the results stand, the constant's tenths digit is 7, and AI systems can contribute mathematically novel steps to long-horizon research while still being unreliable at research judgement and record-keeping.","feed_headline":"New bounds pin Grothendieck constant's tenths digit at 7","feed_subtitle":"A human-AI run proves 6π/11 ≤ K_G ≤ 1.7818, the first lower bound not built from a worst-case example.","key_machinery":"The load-bearing object is the normalized correlation function $H(t)=\\frac{\\pi}{2}\\mathbb{E}[f(X)g(Y)]$ of a Krivine scheme, a pair of odd sign functions on $\\mathbb{R}^k$ used to round SDP vectors to signs via correlated Gaussians. Its Hermite coefficients $b_1,b_3,\\ldots$ control the scheme's rounding quality, and the paper's key identity is the universal affine constraint $b_3\\ge 2b_1-\\frac{11}{6}$, which is linear in the scheme and therefore passes to mixtures and limits. The Naor-Regev theorem then converts any ceiling of this type into a lower bound on $K_G$. For the upper bound, the cubic-quintic limiting scheme works by making the correlation function closer to a straight line in a higher-dimensional limit, suppressing the nonlinear terms that limit half-space rounding.","core_discovery":"The central mathematical claims are two. First, the correlation function $H(t)=b_1t+b_3t^3+\\cdots$ of every Krivine scheme satisfies $b_3\\ge 2b_1-\\frac{11}{6}$; because this affine inequality survives averaging and limits, the Naor-Regev optimality theorem converts it into $K_G\\ge \\frac{6\\pi}{11}$. Second, an explicit cubic-quintic limiting Krivine scheme, obtained by letting the rounding dimension grow, shows $K_G\\le \\frac{\\pi}{2\\log(1+\\sqrt{2})}-3.47\\times10^{-4}$. The paper presents the discovery narrative: the lower bound was found and first proved by the AI system after human operators redirected it from a plateaued upper-bound search, and was then independently verified by the authors.","pith_inferences":["Editorial inference: the same template--prove an affine constraint on the leading coefficients of every rounding scheme, then dualize through an optimality theorem--may transfer to other constants whose optimal rounding families are known, yielding lower bounds from failed upper-bound searches.","Editorial inference: the paper's diagnosis predicts that training models on process records, such as failed approaches, pivots, and lost caveats, rather than finished proofs should specifically improve their research judgement and memory; this could be tested on long-horizon research logs.","Editorial inference: the machine-verified upper bounds $1.781801841033$ and $1.7813319810625639$ were not human-certified; independently checking their certificates would either tighten the interval further or expose a flaw without requiring new mathematics.","Editorial inference: one natural extension is to vary the degree of the polynomial boundary beyond cubic-quintic in limiting Krivine schemes and search numerically for larger inverse-majorant parameters, since the paper's machinery appears designed for systematic expansion of the scheme family."],"forward_implications":["The interval $\\frac{6\\pi}{11}\\le K_G\\le \\frac{\\pi}{2\\log(1+\\sqrt{2})}-3.47\\times10^{-4}$ determines the tenths digit of $K_G$ as 7, replacing the previous wide interval $1.6769\\ldots\\le K_G\\le 1.7822\\ldots$.","The lower-bound proof is the first for $K_G$ that does not construct a hard instance; it establishes a universal ceiling on every Krivine scheme, so future lower-bound work in this framework must respect the constraint $b_3\\ge 2b_1-\\frac{11}{6}$.","The cubic-quintic upper bound answers the question of whether higher-dimensional rounding schemes help, giving the first explicit improvement over Krivine's bound obtained by letting the dimension grow.","The paper's record indicates that a long-horizon AI research system steered by roughly forty human directives can execute novel proof construction but requires human correction of research judgement and record-keeping, as shown by the session-18 pivot and the session-44 withdrawal."],"supporting_citations":[{"why":"Companion paper containing the full proofs of both bounds, including the $b_3\\ge 2b_1-11/6$ theorem and the cubic-quintic certificate.","marker":"[SLX+26]"},{"why":"Proves that Krivine schemes are asymptotically optimal, the transfer theorem that converts a universal ceiling on schemes into $K_G\\ge 6\\pi/11$.","marker":"[NR14]"},{"why":"Gives the classical upper bound $\\pi/(2\\log(1+\\sqrt{2}))$ and the arcsine identity for half-space rounding, the baseline that both new bounds improve.","marker":"[Kri77]"},{"why":"Disproves Krivine's conjecture and asks whether higher-dimensional schemes help, the question answered affirmatively by the limiting cubic-quintic scheme.","marker":"[BMMN11]"},{"why":"Original Grothendieck inequality that defines the constant and frames the SDP integrality-gap question.","marker":"[Gro53]"},{"why":"The previous lower bound $1.6769\\ldots$ via Gaussian construction that the new lower bound improves, and whose non-constructive mechanism it replaces.","marker":"[Dav84, Ree91]"}],"fun_headline_variants":["AI tightens Grothendieck constant bounds","New lower bound for Grothendieck constant from AI","AI finds tighter Grothendieck constant lower bound","Human-AI collaboration yields better K_G bounds","Grothendieck constant bound improved by AI proof"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The lower bound rests on the assumption that an inequality verified for finitely many rounding schemes survives averaging and limits; the paper records that this transfer was the part repaired during adversarial review, and if the repair is wrong the bound $K_G\\ge 6\\pi/11$ does not follow.","fun_headline_variants_meta":{"raw":{"variants":["AI tightens Grothendieck constant bounds","New lower bound for Grothendieck constant from AI","AI finds tighter Grothendieck constant lower bound","Human-AI collaboration yields better K_G bounds","Grothendieck constant bound improved by AI proof"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00057,"raw_usage":{"total_tokens":2679,"prompt_tokens":909,"completion_tokens":1770,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1695}},"tokens_in":525,"tokens_out":1770,"duration_ms":14576,"temperature":1.0,"reasoning_tokens":1695,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:17:12.681764+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently recompute, in interval-certified arithmetic, the one-dimensional Gaussian inequalities used in the companion paper's proof that every Krivine scheme satisfies $b_3\\ge 2b_1-11/6$; a single certified violation for a mixed or limiting scheme would refute the lower-bound proof.","supporting_citations":[],"review_version":1}