{"id":"65b4296c-7d8c-469f-af05-126e3c25663a","arxiv_id":"2411.12004","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"High-level CCSDT and CCSDT(Q) calculations on the S66 benchmark show that the error cancellation in CCSD(T) breaks down for pi-stacking complexes, which are overbound, and simple fitted formulas estimate the correction to about 0.01 kcal/mol.","lead":"High-level quantum chemistry calculations on 66 molecular dimers show that the standard CCSD(T) method slightly overbinds pi-stacked pairs, while other interaction types stay well balanced. The paper also offers cheap formulas to estimate this correction, so larger biomolecular complexes can be re-evaluated without prohibitively expensive calculations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pi-stack overbinding magnitude rests on CCSDT(Q)-CCSD(T) corrections computed only in a double-zeta basis; a larger-basis check could shift or even overturn the 0.04-0.10 kcal/mol estimate.","rationale":"The reader's weakest_assumption and my concern coincide: the small double-zeta basis for post-CCSD(T) corrections is the main correctness risk. I do not see an internal inconsistency or a fatal flaw. The paper is transparent about the basis, provides raw data, and gives genuine out-of-sample checks (naphthalene dimer, S22 subset) that support the sign and rough size of the pi-stacking effect. The residual concern is that the quantitative range, and especially the overbinding magnitude emphasized in the abstract and conclusions, might depend on basis-set effects that have not yet been measured directly. A single larger-basis CCSDT(Q) datapoint on a representative pi-stack would substantially settle this. Since the reader already conditioned the verdict on this issue, my assessment does not change the conditional verdict.","tokens_in":16225,"tokens_out":6879,"duration_ms":72661,"concrete_test":"Perform a rank-reduced or reduced-scaling CCSDT(Q) calculation (e.g., Lesiuk's implementation, as used for naphthalene dimer in Ref. [22]) for benzene...pyridine (system 27) or pyridine...pyridine (system 25) in the cc-pVTZ or haVTZ(f,d) basis, and compare the CCSDT(Q)-CCSD(T) interaction-energy correction with the cc-pVDZ(d,s) value in Table 1/4. If the correction shifts by more than roughly 0.03 kcal/mol or changes sign, the DZP-based overbinding estimate is not representative. A less expensive complementary check is CCSDT-CCSD(T) in cc-pVTZ for the same systems, since higher-order triples carry most of the effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central numerical result for pi-stacking overbinding relies entirely on CCSDT(Q)-CCSD(T) corrections computed in the cc-pVDZ(d,s) basis (Section 2, Table 2). The only larger-basis sanity check, the naphthalene dimer comparison (Section 3), uses cc-pVDZ(d,p), still a double-zeta basis. Table 3 shows that the ratio Delta E(T)/Delta Ecorr[CCSD] used by Eq. (12) is strongly basis-set dependent: the fitted beta changes from -0.273 at cc-pVDZ(d,s) to -1.009 at haVTZ(f,d). This does not by itself prove that the actual T3-(T) and (Q) interaction-energy corrections are basis-dependent, but it removes any default expectation that DZP post-CCSD(T) corrections are converged. If those corrections shift by more than roughly 0.03 kcal/mol at larger basis, the claimed 0.04-0.10 kcal/mol overbinding range (Table 4) would need revision, and the category-level message would rest on the sign alone. The FN-DMC comparison is too noisy to arbitrate, and the note added in revision reports an FN-DMC S66 result for hydrogen bonds that is hard to rationalize, further weakening the external validation. The concern is a correctness risk, not an internal inconsistency: the authors are transparent about the basis, but the headline physical claim assumes transferability to the CCSD(T)/CBS reference values.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports CCSDT and CCSDT(Q) calculations (where feasible) for the S66 noncovalent interaction benchmark in the cc-pVDZ(d,s) basis, and analyzes post-CCSD(T) contributions decomposed by interaction category. The central finding is that for hydrogen-bonded, pure London, and mixed-influence complexes, the error cancellation between higher-order triples (T3-(T)) and connected quadruples (Q) makes CCSD(T) close to CCSDT(Q), while for pi-stacked systems this cancellation breaks down and CCSD(T) overbinds. The authors also propose simple fitted formulas, Eqs. (12) and (13), that reproduce CCSDT(Q)-CCSD(T) differences to about 0.01 kcal/mol RMS on the training set, and they compare with CCSD(T)_Lambda, FN-DMC results, a naphthalene dimer check, and a W4-11 thermochemical application.","tokens_in":16609,"tokens_out":9396,"duration_ms":84815,"significance":"If the double-zeta results are representative of the basis-set limit, the paper provides valuable direct evidence about the breakdown of error cancellation in CCSD(T) for pi-stacked complexes, a question of current interest in benchmarking noncovalent interactions. The calculations are state-of-the-art in scope, the authors are transparent about missing data, and the naphthalene dimer and S22 checks are genuine out-of-sample tests. The proposed low-scaling estimators are useful in practice, and the W4-11 semiquantitative transferability test is a positive feature. The significance is conditional, however, on two points: the basis-set transferability of the post-CCSD(T) corrections, and the distinction between actual CCSDT(Q) data and model estimates in the category-level statistics.","major_comments":[{"comment":"The central quantitative claim that CCSD(T) overbinds pi-stacks by about 0.04-0.10 kcal/mol rests entirely on post-CCSD(T) corrections computed in the cc-pVDZ(d,s) basis. The only larger-basis comparison, the naphthalene dimer check, uses cc-pVDZ(d,p), which is still a double-zeta basis. Table 3 shows that the ratio Delta_E(T)/Delta_E_corr[CCSD] underlying Eq. (12) is strongly basis-set dependent, with beta changing from -0.273 at cc-pVDZ(d,s) to -1.009 at haVTZ(f,d). This does not by itself prove that the actual T3-(T) and (Q) interaction-energy corrections are basis-dependent, but it removes any default expectation that DZP post-CCSD(T) corrections are converged. Please provide at least a few larger-basis CCSDT(Q) or CCSDT(Q)-level calculations for representative pi-stacks, or state clearly that the quantitative overbinding range is a DZP result and that the robust conclusion is the sign, not the magnitude.","section":"Section 3 (Tables 2 and 4), with Table 3"},{"comment":"The headline RMSD of about 0.01 kcal/mol for Eq. (12) is an in-sample fit: the parameters alpha, beta, and a3 are optimized on S66 and then evaluated on the same S66 data. The out-of-sample evidence is considerably weaker: the S22 subset gives RMSD 0.021 kcal/mol, which the authors attribute to the formic acid dimer outlier, and the naphthalene dimer estimates (-0.150 and -0.201 kcal/mol for Eq. (12) and Eq. (13), respectively) bracket the actual value (-0.160 kcal/mol) with a spread comparable to the effect itself. I therefore cannot parse the abstract claim that the model 'predicts CCSDT(Q)-CCSD(T) differences to 0.01 kcal/mol RMS' as a predictive statement. Please state explicitly that the 0.01-kcal/mol figure is the training-set error, and give the out-of-sample RMS excluding the identified outlier.","section":"Section 3, Eq. (12) and Table 3"},{"comment":"It appears that actual CCSDT(Q)-CCSD(T) values are available for only two pi-stacking systems, 24 and 25, as seen in the 'actual' column of Table 4, while the entries for systems 26-29 and 47-49 are model estimates shown in parentheses. Yet Table 2 reports category averages, such as a pi-stack average of -0.043 kcal/mol for CCSDT(Q)-CCSD(T), without indicating which entries are actual and which are estimated. The statement in the conclusions that 'one observes nontrivial discrepancies between CCSD(T) and CCSDT(Q)' for small aromatic pi-stacks is therefore based on a very small number of direct calculations augmented by the fitted models. Please specify the number of actual CCSDT(Q) data per category and, if estimates are used, provide separate statistics for actual and estimated entries.","section":"Section 3 (Table 4) and Table 2"}],"minor_comments":[{"comment":"The meaning of parentheses in Table 4 (model estimates) is explained only in the caption; please state this explicitly in the table itself or in the text, and mark the 'actual' column entries that are genuine CCSDT(Q) data.","section":"Section 3, Table 4"},{"comment":"The abstract and highlights say the present results 'corroborate' FN-DMC claims, but the FN-DMC uncertainties in Table 4 are large, and the note added in revision reports FN-DMC S66 hydrogen-bond discrepancies that are hard to rationalize. Consider softening the wording to 'consistent in sign with FN-DMC for pi-stacks' to avoid overstating the external validation.","section":"Abstract and Highlights"},{"comment":"The notation 'cc-pVDZ(p,s)' appears once in Section 3 in the context of scaling Delta_E[(Q)]; please define it in Section 2 alongside cc-pVDZ(d,s) and cc-pVDZ(d,p).","section":"Section 2, basis set notation"},{"comment":"Eq. (3) is attributed to Grueneis et al. [15], but the parameters a1 and a2 are fitted in this work; please clarify explicitly which parameters are from Ref. [15] and which are refitted here.","section":"Section 3, Eq. (3)"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for Chemical Physics Letters and the central physical finding is plausible and interesting. My major concerns are the basis-set transferability of the post-CCSD(T) corrections and the transparency of the actual-versus-estimated data underlying the category statistics; both are fixable with additional calculations or careful rewording. I do not see grounds for rejection, but the current manuscript overstates the quantitative certainty of the pi-stack overbinding range and the predictive power of the fitted formulas."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read Semidalas, Boese, Martin on post-CCSD(T) corrections in S66. The main deliverable is a new dataset: CCSDT and CCSDT(Q) corrections for 59 of the 66 S66 complexes in a DZP basis. That's a real step beyond prior work, which covered only a handful of dimers or the nine pi-stacks from the Grüneis group. The physical message is clear and likely right: error cancellation between higher-order triples and connected quadruples holds for most categories but breaks down for aromatic pi-stacking, where CCSD(T) overbinds. The direction is corroborated by the naphthalene dimer check and by independent FN-DMC/CCSD(cT) comparisons, even though those are noisy.\n\nThe paper does several things well. The perturbation-theory decomposition of the contributions is instructive, and the empirical models are honest: they report in-sample RMSD ~0.01 kcal/mol but then show out-of-sample checks on S22 and W4-11 where accuracy degrades, with an identified outlier. That kind of transparency deserves credit.\n\nThe soft spots are real but not fatal. The headline corrections are computed only in cc-pVDZ(d,s). The authors know this and even report that the fitted parameter beta in Eq. (12) changes from -0.273 at DZP to -1.009 at haVTZ(f,d), which is a strong hint that the model parameters are basis-sensitive. They don't test whether the actual T3-(T) and (Q) interaction-energy corrections are similarly basis-dependent; the naphthalene dimer check is still double-zeta. So the magnitude of the pi-stack overbinding, 0.04-0.10 kcal/mol, should be treated as a DZP estimate rather than a CBS number. The direction is probably safe, but a larger-basis calculation on one or two pi-stacks would solidify it. The FN-DMC comparison is noisy, and the note added in revision about FN-DMC overbinding H-bonds by 0.7 kcal/mol makes FN-DMC a shaky arbiter, not the authors' fault but it dims the external validation.\n\nThe RMSD of the fitted models is in-sample, but they don't oversell it; they show genuine out-of-sample degradation. This is a minor issue, not a load-bearing flaw.\n\nWho should read this: anyone building benchmark references for noncovalent interactions, and method developers working on post-CCSD(T) corrections or DMC comparisons. It deserves a serious referee. My recommendation: send it to review, with a request that the authors either add a larger-basis sanity check for a pi-stack or explicitly qualify the magnitude claim as basis-dependent.","headline":"Valuable new post-CCSD(T) dataset for S66; pi-stack overbinding direction is likely right but its magnitude rests on a DZP basis and needs a larger-basis check.","tokens_in":17134,"tokens_out":4089,"would_cite":true,"duration_ms":36581,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CCSD(T) overbinds pi-stacked complexes because the usual cancellation between higher-order triples and connected quadruples fails, by about 0.04-0.10 kcal/mol in the S66 benchmark.","keywords":["noncovalent interactions","S66 benchmark","coupled cluster theory","CCSD(T)","post-CCSD(T) corrections","pi-stacking","higher-order triples","connected quadruples"],"falsifier":"Compute CCSDT(Q)-CCSD(T) for one of the parallel-displaced aromatic stacks, such as the benzene dimer, in a triple-zeta basis set; if the correction is substantially smaller in magnitude than -0.05 kcal/mol, the small-basis overbinding estimate does not survive basis-set extrapolation. Alternatively, run fixed-node diffusion Monte Carlo for systems 24-29 with stochastic error below 0.02 kcal/mol and compare with the CCSD(T)/CBS values.","tokens_in":16025,"feed_emoji":"🧪","tokens_out":19678,"duration_ms":147117,"temperature":0.7,"pith_summary":"Using the S66 benchmark of biomolecular dimers, the paper computes full CCSDT and CCSDT(Q) corrections in a polarized double-zeta basis to test the assumption that CCSD(T) is essentially exact for noncovalent interactions. It finds that for hydrogen bonds, pure London complexes, and mixed-influence systems, the usual error cancellation holds: repulsive higher-order triples ($T_3-(T)$) are offset by attractive connected quadruples $(Q)$. For $\\pi$-stacking complexes this cancellation starts to fail, leaving CCSD(T) overbound: the average CCSDT(Q)-CCSD(T) correction for the $\\pi$-stack subset is -0.043 kcal/mol, with the fully converged parallel-displaced dimers at about -0.085 and -0.099 kcal/mol and model estimates for the remaining stacks reaching near -0.11 kcal/mol. The paper also shows that CCSD(T)$_\\Lambda$ fixes the $\\pi$-stack problem but overbinds London complexes, and that a two-parameter formula estimates the full CCSDT(Q)-CCSD(T) difference to 0.01 kcal/mol RMS at no steeper than $O(N^7)$ cost.","feed_headline":"CCSD(T) overbinds pi-stacked complexes by up to 0.1 kcal/mol","feed_subtitle":"A two-parameter formula estimates the missing higher-order correlation at O(N^7) cost, keeping benchmarks tractable.","key_machinery":"The load-bearing objects are the fifth-order perturbation terms of Eq. (7): the connected-quadruples self-interaction $E^{(5)}_{QQ}$, the triples-quadruples coupling $E^{(5)}_{TQ}$, and the triples-triples interaction $E^{(5)}_{TT}$; together they make up the CCSDT(Q)-CCSD(T) difference through fifth order. The paper identifies $E^{(5)}_{QQ}$ as almost universally attractive and $E^{(5)}_{TQ}$ as antibonding, with the triples-triples term distinguishing $\\pi$-stacks (antibonding) from London complexes (slightly bonding). The second key object is the two-parameter estimator of Eq. (12), $\\Delta E[\\mathrm{postCCSD(T)}] \\approx \\Delta E(T)[\\alpha/(1-\\beta\\, \\Delta E(T)/\\Delta E_{\\mathrm{corr}}[\\mathrm{CCSD}]) - 1] + a_3(\\Delta E[\\mathrm{CCSDT-3}] - \\Delta E[\\mathrm{CCSD(T)}])$, with $a_3=1$; it converts the cheaply accessible $(T)$ energy and CCSDT-3 correction into a prediction of the full CCSDT(Q) correction. CCSDT-3 appears because it captures the lion's share of the higher-order triples effect at $O(n^3_{\\rm occ}N^4_{\\rm virt})$ scaling, and the geometric-series form of the first term mimics the convergence of the correlation-energy series.","core_discovery":"The paper's central claim is that CCSD(T), long treated as the gold standard for noncovalent interactions, is not uniformly reliable across the S66 dataset: for $\\pi$-stacked aromatic dimers the cancellation between the (usually repulsive) higher-order triples correction $T_3-(T)$ and the (attractive) connected quadruples correction $(Q)$ breaks down. Higher-order triples become strongly repulsive for these stacks, and while $(Q)$ is also largest there, it does not fully compensate, yielding net CCSDT(Q)-CCSD(T) corrections of about -0.04 kcal/mol on average for the $\\pi$-stack subset, with fully converged parallel-displaced dimers at -0.085 and -0.099 kcal/mol and model estimates for the remaining stacks reaching about -0.109 kcal/mol. The authors show by a fifth-order analysis, $E[\\mathrm{CCSDT(Q)}]-E[\\mathrm{CCSD(T)}] = E^{(5)}_{QQ} + 2E^{(5)}_{TQ} + E^{(5)}_{TT} + O(\\lambda^6)$, that the residual is driven by the triples-triples and triples-quadruples interaction terms. They further demonstrate that CCSD(T)$_\\Lambda$ moves the error from $\\pi$-stacks to London complexes, and that a two-parameter model based on $(T)$, the CCSD correlation energy, and the CCSDT-3 correction reproduces the post-CCSD(T) difference with 0.01 kcal/mol RMS, allowing cheap $O(N^7)$ estimates of what full CCSDT(Q) would give.","pith_inferences":["If the overbinding grows with stack size as the naphthalene dimer check suggests, then CCSD(T)-based reference values for larger aromatic $\\pi$-stacks, such as DNA base-pair steps or organic semiconductor dimers, may carry a systematic bias that is currently unaccounted for in benchmark tables.","The strong basis-set dependence of $\\beta$ in Eq. (12) implies that the formula should be re-fitted rather than used with S66 parameters when applied to thermochemistry or to complexes with different monomer polarizabilities.","A natural testable extension is to apply Eq. (12) to a new set of $\\pi$-stacked dimers with experimental binding energies; agreement within the claimed 0.01 kcal/mol RMS would support using it as a correction scheme for cheaper methods like DFT or MP2."],"forward_implications":["For aromatic $\\pi$-stacking complexes in S66, CCSD(T)/CBS binding energies are overbound by roughly 0.04-0.10 kcal/mol, so reference tables built on CCSD(T) for such systems carry this bias.","CCSD(T)$_\\Lambda$ removes most of the $\\pi$-stack overbinding but overshoots for pure London complexes, so neither CCSD(T) nor CCSD(T)$_\\Lambda$ alone matches CCSDT(Q) across all interaction categories.","The two-parameter model of Eq. (12) predicts CCSDT(Q)-CCSD(T) to 0.01 kcal/mol RMS for S66 and brackets the naphthalene dimer value, giving a practical $O(N^7)$ proxy for the full correction.","CCSDT-3 captures most of the higher-order triples effect and is a reliable stand-in for full CCSDT when estimating post-CCSD(T) corrections.","The gap between CCSD(T) and CCSDT(Q) for $\\pi$-stacks is expected to widen for larger aromatic stacks."],"supporting_citations":[{"why":"Supplies the S66 benchmark systems and reference geometries that define the test set.","marker":"[28]"},{"why":"Defines the CCSD(T) method whose accuracy and error compensation are being assessed.","marker":"[5]"},{"why":"Defines full CCSDT, needed to isolate the higher-order triples contribution $T_3-(T)$.","marker":"[31]"},{"why":"Defines CCSDT(Q), the reference including connected quadruples used for the post-CCSD(T) corrections.","marker":"[33]"},{"why":"Provides the CCSD(cT) results and the claimed discrepancies for $\\pi$-stacks that motivate the study and give comparison data.","marker":"[15]"},{"why":"Supplies fixed-node diffusion Monte Carlo reference values for large $\\pi$-stacking complexes used to benchmark CCSD(T).","marker":"[11]"},{"why":"Provides reduced-scaling CCSDT(Q) values for naphthalene and benzene sandwiches used as a sanity check for the fitted models.","marker":"[22]"},{"why":"Supports neglecting basis set superposition error in post-CCSD(T) corrections.","marker":"[40]"},{"why":"Defines the CCSDT-n approximations, including CCSDT-3 and CCSDT-1b, that underlie the cheap estimators.","marker":"[17]"}],"fun_headline_variants":["CCSD(T) overbinds pi-stacked dimers; cheap model fixes it","Pi-stacked complexes break CCSD(T) accuracy; new O(N^7) model corrects","CCSD(T) error for pi-stacks up to 0.1 kcal/mol, model predicts","Pi-stacking breaks CCSD(T) error cancellation; cheap model restores accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The overbinding estimate rests on the assumption that post-CCSD(T) corrections computed in the small cc-pVDZ(d,s) basis set represent the complete basis set limit, even though the model's $\\beta$ parameter changes from -0.273 to -1.009 when the basis is enlarged.","fun_headline_variants_meta":{"raw":{"variants":["CCSD(T) overbinds pi-stacked dimers; cheap model fixes it","Pi-stacked complexes break CCSD(T) accuracy; new O(N^7) model corrects","CCSD(T) error for pi-stacks up to 0.1 kcal/mol, model predicts","Pi-stacking breaks CCSD(T) error cancellation; cheap model restores accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":2987,"prompt_tokens":1066,"completion_tokens":1921,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":682,"completion_tokens_details":{"reasoning_tokens":1825}},"tokens_in":682,"tokens_out":1921,"duration_ms":14504,"temperature":1.0,"reasoning_tokens":1825,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:01:03.821548+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute CCSDT(Q)-CCSD(T) for one of the parallel-displaced aromatic stacks, such as the benzene dimer, in a triple-zeta basis set; if the correction is substantially smaller in magnitude than -0.05 kcal/mol, the small-basis overbinding estimate does not survive basis-set extrapolation. Alternatively, run fixed-node diffusion Monte Carlo for systems 24-29 with stochastic error below 0.02 kcal/mol and compare with the CCSD(T)/CBS values.","supporting_citations":[{"cited_title":"Řezáč, K.E","cited_arxiv_id":null,"evidence_quote":"Supplies the S66 benchmark systems and reference geometries that define the test set."},{"cited_title":"Raghavachari, G.W","cited_arxiv_id":null,"evidence_quote":"Defines the CCSD(T) method whose accuracy and error compensation are being assessed."},{"cited_title":"Noga, R.J","cited_arxiv_id":null,"evidence_quote":"Defines full CCSDT, needed to isolate the higher-order triples contribution $T_3-(T)$."},{"cited_title":"Bomble, J.F","cited_arxiv_id":null,"evidence_quote":"Defines CCSDT(Q), the reference including connected quadruples used for the post-CCSD(T) corrections."},{"cited_title":"Al-Hamdani, P.R","cited_arxiv_id":null,"evidence_quote":"Supplies fixed-node diffusion Monte Carlo reference values for large $\\pi$-stacking complexes used to benchmark CCSD(T)."},{"cited_title":"Fishman, E","cited_arxiv_id":null,"evidence_quote":"Supports neglecting basis set superposition error in post-CCSD(T) corrections."},{"cited_title":"Noga, R.J","cited_arxiv_id":null,"evidence_quote":"Defines the CCSDT-n approximations, including CCSDT-3 and CCSDT-1b, that underlie the cheap estimators."}],"review_version":1}