{"id":"3115a0ab-95ba-402f-9719-95c2e02df3da","arxiv_id":"1908.01054","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A held-out particle set and BioEM posterior calculations can distinguish overfitted cryo-EM reconstructions from well-behaved ones by checking whether map probability grows with added high-frequency detail.","lead":"The authors test a cross-validation method for cryo-EM density maps that holds out a small set of particle images and checks whether the probability of the map given those images grows as higher-resolution frequencies are added. On three standard maps the probability rises as expected, and on two suspicious datasets it does not, suggesting the tests can flag overfitted reconstructions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The control-set posterior is not a marginal held-out likelihood: round-1 orientation selection uses the final map under validation, so the kc-trend that discriminates overfitting may reflect orientation overfitting of control particles.","rationale":"The reader identified the same weakest assumption, and I agree: the orientation search using the final reconstruction is the place where the 'independent control set' claim is least secure. If per-particle orientations are optimized, even locally, against the control images using the map under validation, the control set is no longer a clean held-out sample; this directly threatens the conceptual foundation of the paper, not just the numerical demonstration. The paper's Eq. (2) suggests a proper marginal, but the described double-round algorithm selects best orientations rather than integrating over the full orientation space, so the computed posterior is an upper-biased approximation whose bias can grow with kc for noise-like maps. The fact that the overfitted examples in Fig. 2 are a single controversial map and a synthetic noise control makes the demonstration insufficient to dispel this confound. A targeted recomputation can settle whether the kc-trends survive an orientation search that does not use the final map. The paper is otherwise careful and the proposed protocol is useful; with this test, I would keep the reader's conditional recommendation rather than reject or accept outright.","tokens_in":12592,"tokens_out":8729,"duration_ms":91541,"concrete_test":"Recompute the BioEM evaluation for the control sets of HIV-ET and the RAG1-RAG2 noise control using a uniform orientation integration over the full 36864-quaternion grid in round 1 (no selection based on the final reconstruction), followed by the same round-2 protocol; alternatively, use a 60-Å low-pass-filtered early-iteration map for round-1 orientation selection. If the cumulative log-posterior vs kc for the overfitted sets becomes increasing, or the separation from the standard maps collapses, the discrimination is an orientation-fitting artifact. If the trends are unchanged, the concern is mitigated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (abstract: 'able to discriminate the overfitted sets from the non-overfitted ones') rests on the cumulative log-posterior over an independent control set increasing with frequency cutoff for genuine maps and not for overfitted ones. The independence of this statistic is compromised by the BioEM orientation search. In Methods ('BioEM algorithm'), round 1 obtains 'the best orientations for each particle' using 'the final reconstruction from the refinement ... without low-pass filtering'—i.e., the very map being validated. Round 2 then re-fits orientations for each low-pass filtered map, zooming around those 10 best orientations with 1250 quaternions. Equation (2) presents the posterior as an integral over Θ, but the implementation is a mode-based search, not a full marginalization over orientations with a fixed prior. This turns the held-out posterior into a profile likelihood in which per-particle nuisance parameters are optimized against the control set. For an overfitted map, whose high-frequency content is noise, adding frequency shells provides more features that a per-particle orientation search can use to match noise in the control images. The observed increase in log-posterior with kc could therefore be an artifact of this orientation fitting, not genuine generalization. The discrimination is tested on only one real overfitted case (HIV-ET) plus a synthetic pure-noise control, so this confound has not been ruled out.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a cross-validation protocol for cryo-EM reconstructions based on a small set of particle images excluded from refinement. For each gold-standard reconstruction, the authors low-pass filter the map at several cutoffs, compute the BioEM posterior probability of each filtered map over the control particles, and monitor the cumulative log-posterior and a normalized Jensen-Shannon divergence (NJSD) between the two half-set reconstructions as functions of frequency cutoff and refinement iteration. They report that for three standard maps (HCN1, TRPV1, RAG1-RAG2) the cumulative log-posterior increases with cutoff and iteration and the NJSD follows an inverse-exponential trend whose scale gamma correlates with the FSC-based resolution, whereas for an HIV-1 envelope trimer map and a synthetic pure-noise control the expected trends fail. The authors conclude that an independent control set is essential for unbiased map validation.","tokens_in":12912,"tokens_out":6788,"duration_ms":62727,"significance":"The proposed tool is a welcome addition to the cryo-EM validation toolbox: it uses raw particle data rather than post-processed maps, is conceptually analogous to R-free, and the authors provide open-source code, a tutorial, and tests on public EMPIAR datasets. If the reported trends are reproducible with truly held-out orientations, the method could provide a practical overfitting diagnostic complementary to the gold-standard FSC. However, the current manuscript's central claim is weakened by a statistical independence gap in the orientation fitting and by the limited evidence for the overfitted category.","major_comments":[{"comment":"The cumulative log-posterior reported in Fig. 2 is not a marginal held-out likelihood. In round 1 of the BioEM calculation, the best 10 orientations for each control particle are selected using 'the final reconstruction from the refinement with a broad mask and without low-pass filtering' (Methods, BioEM algorithm), i.e., the map being validated; in round 2, orientations are re-fitted for each low-pass filtered map by zooming around those 10 orientations. Thus the posterior in Eq. (2), which integrates over orientations, is approximated by a profile likelihood in which per-particle orientations are optimized against the control set. For an overfitted map whose high-frequency content is noise, the orientation search can align that noise with the control images, so the observed increase of cumulative log-posterior with kc (or its absence for overfitted maps) may reflect orientation overfitting rather than genuine generalization. Please show that the trends in Figs. 2 and 3 persist when orientations are fixed across all kc, for example by using the orientations produced by the gold-standard RELION refinement of the non-control particles, or by actually marginalizing over orientations.","section":"Methods: BioEM algorithm"},{"comment":"The evidence for the central discrimination claim rests on only one real overfitted dataset (HIV-ET) and one synthetic pure-noise control. The HIV-ET map is externally controversial (refs. [42,43]), and the pure-noise control is not a reconstruction at all but a set of unrelated Gaussian images evaluated against RAG1-RAG2 maps; its failure is therefore an expected consequence of mismatched data rather than a demonstration of overfitting detection. To support the claim that the method 'discriminates the overfitted sets from the non-overfitted ones,' the authors should include additional known overfitted reconstructions (e.g., maps obtained by refining noise-substituted data or other disputed EMPIAR deposits) and report results over multiple random selections of the control set, with error bars.","section":"Results: Map evidence from the cumulative log-posterior"},{"comment":"The correlation between the fitted frequency gamma and the inverse FSC resolution (Fig. 4, r²=0.93, 0.91, 0.85) is reported without uncertainties or goodness-of-fit measures. Each NJSD curve is fitted with three free parameters (A, B, gamma) to only eight cutoff frequencies, so the statistical significance of these correlations is unclear. Please report parameter errors, per-system fits, and the sensitivity of gamma to the fitting model.","section":"Results: Cross-validation tests versus resolution"}],"minor_comments":[{"comment":"The NJSD is defined with probabilities normalized per image such that P1ω + P2ω = 1; this is stated in the text but not in the equation, so consider making the normalization explicit in Eq. (3) to avoid ambiguity.","section":"Methods: Eq. (3)"},{"comment":"The gradient color scale from maroon to green over refinement iterations is difficult to read without a legend; adding iteration numbers or a color bar would improve interpretability.","section":"Figure 2"},{"comment":"The caption contains a typo: 'NSJD' should be 'NJSD'.","section":"Supplementary Figure 3 caption"},{"comment":"The claim that the method 'converge[s] over a small particle set, typically only 1000 particles' is supported by a single system (TRPV1) at a single iteration; please either weaken the claim or show convergence for more systems.","section":"Discussion"},{"comment":"The low-pass filter is implemented as a hard cutoff in Fourier space (Eq. 1); a soft-edged filter or a Butterworth filter might be preferable to avoid Gibbs ringing, and the choice deserves a sentence of justification.","section":"Methods: Low-pass filter"}],"recommendation":"major_revision","confidential_remarks":"The main risk to publication is the orientation-search confound. If the authors can demonstrate the trends are robust to fixing orientations (or provide a valid marginalization over orientations), the paper would be a useful contribution. The HIV-ET case being controversial and the synthetic noise control being trivially mismatched also need stronger evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper transplants the R-free concept to cryo-EM: hold out a small particle set, then monitor the BioEM posterior of the reconstructions over that set as a function of frequency cutoff. The specific two-test design - cumulative log-posterior and a normalized Jensen-Shannon divergence between half-map distributions - is new, and the demonstration on three public datasets plus two overfit cases is genuinely useful. The expected trends appear for the genuine maps and break down for HIV-ET and the pure-noise control, suggesting the method has discriminating power. They also show convergence with about 1000 particles and ship code and tutorials, which is real value.\n\nThe main soft spot is the one the stress-test flags, and it is real. Round 1 of the BioEM orientation search uses the final unfiltered reconstruction - the very map being validated - to pick the best orientations for each control particle. So the control set is not independent in the strict sense; the resulting statistic is a profile likelihood with per-particle orientation parameters fit to the map. That could inflate the posterior for overfit maps whose high-frequency noise is matched by the orientation search, weakening the claimed discrimination. The fact that the discrimination still works on the tested cases is encouraging, but one controversial real map (HIV-ET) and a synthetic noise set are not enough to rule out the confound. No error bars or repeated control sets either, so we do not know how stable the trends are.\n\nA smaller quibble: the gamma-versus-resolution correlation is an empirical fit with only three systems, and the inverse-exponential shape is not derived from the physics. That is fine as a heuristic, but the paper should be more careful about claiming it as a resolution estimator.\n\nOverall, the central idea is worth taking seriously, but the current evidence is conditional. A serious referee should ask for either a proper marginalization over orientations with a fixed prior, or a control experiment on simulated data with known overfitting to show the round-1 orientation choice is not driving the signal. I would definitely send this out for review rather than desk-reject, and I would bring it to a reading group to discuss the orientation conundrum.","headline":"A plausible R-free-like control-set test for cryo-EM maps that deserves a serious referee, but the held-out set is not fully independent because the orientation search uses the map under validation.","tokens_in":689,"tokens_out":1570,"would_cite":true,"duration_ms":53282,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cryo-EM reconstructions can be cross-validated by watching how their Bayesian posterior behaves over a small withheld particle set.","keywords":["cryo-EM","map validation","cross-validation","control particle set","BioEM posterior","overfitting","normalized Jensen-Shannon divergence","gold-standard refinement"],"falsifier":"Take a deliberately overfitted map and rerun the cross-validation with the control particles' orientations fixed by a reference map unrelated to the refinement, or randomized; if the rising cumulative log-posterior and NJSD signature reappears, the discrimination is an artifact of the orientation search rather than a property of genuine maps.","tokens_in":12395,"feed_emoji":"🔬","tokens_out":8537,"duration_ms":83281,"temperature":0.7,"pith_summary":"The paper proposes that cryo-EM reconstructions should be checked against a small control set of particle images withheld from refinement, in the spirit of the R-free test in crystallography. For each gold-standard half-set map, low-pass filtered to several frequency cutoffs, the authors compute a Bayesian posterior probability of the map over the control set. They claim that for genuine maps the cumulative log-posterior increases with the frequency cutoff and with refinement iteration, while for overfitted maps it decreases or stays flat; and that the normalized Jensen-Shannon divergence between the two half-sets' probability distributions grows with frequency only for genuine maps. The authors show the resulting signatures separate three standard cryo-EM systems from two overfitted cases, and that roughly 1000 control particles are enough for the observables to converge. If correct, the protocol supplies a direct raw-data-based map validation that complements the gold-standard FSC.","feed_headline":"A withheld particle set can spot overfitted cryo-EM maps","feed_subtitle":"In genuine maps the map's probability over held-out particles rises with frequency; in overfitted maps it does not.","key_machinery":"The central object is the BioEM posterior probability $P_{m\\omega}$: a Bayesian likelihood of a map $m$ given a single particle image $\\omega$, with nuisance parameters (center displacement, normalization, offset, noise, orientation, CTF defocus, amplitude, B-factor) integrated out. The cumulative log-posterior over the control set, $\\sum_\\omega \\ln P_{m\\omega}/N_\\omega - \\ln P_{\\mathrm{noise}}$, acts as the first test statistic, and the normalized Jensen-Shannon divergence between the posterior distributions of the two half-set reconstructions acts as the second. The integrations are done in two rounds: a coarse all-orientation search selects the best ten orientations per particle, then a fine zoom around those orientations scores each low-pass filtered map. This machinery converts a held-out particle set into a quality ladder: genuine maps climb it as frequencies are added and refinement proceeds, while overfitted maps do not.","core_discovery":"Using the BioEM posterior, a per-image likelihood of a map integrated over orientation, CTF, displacement, and noise parameters, as a map-quality statistic over an independent particle set, the paper shows that the cumulative log-posterior grows with low-pass frequency cutoff and refinement iteration for the genuine reconstructions of HCN1, TRPV1, and RAG1-RAG2, while it decreases or stays constant for a low-resolution HIV-1 envelope trimer reconstruction and for reconstructions evaluated against pure-noise images. The normalized Jensen-Shannon divergence between the probability distributions of the two gold-standard half-set reconstructions rises with added high frequencies and plateaus for the genuine systems, and the fitted plateau frequency $\\gamma$ correlates strongly with the inverse of the FSC resolution ($r^2 = 0.93$, $0.91$, and $0.85$). The authors conclude that withholding a control particle set from refinement makes overfitting detectable and should become standard practice.","pith_inferences":["Because the first round of the posterior calculation selects each control particle's best orientations against the final reconstruction being validated, the control set is not fully independent under the protocol as written; a stricter implementation that fixes orientations from an unrelated reference would test whether the discrimination survives.","The strong correlation between $\\gamma$ and the inverse FSC resolution suggests the method could be calibrated on a collection of deposited maps into an absolute resolution estimator that does not depend on a mask; the paper does not attempt that calibration.","The per-particle probabilities that enter the cumulative log-posterior could be mined for diagnostics: particles where the two half-set reconstructions disagree most are prime candidates for heterogeneity or misassigned defocus, an application the paper does not explore."],"forward_implications":["Every cryo-EM refinement could carry a small control set, giving an R-free-like, raw-data-based check on map quality at each iteration.","The NJSD plateau frequency $\\gamma$ tracks the inverse FSC resolution ($r^2 = 0.93$, $0.91$, $0.85$), so resolution information can be extracted from the control set rather than from mask-dependent FSC alone.","The cross-validation observables converge with roughly 1000 control particles, so the extra cost of the protocol is small relative to a full refinement.","Overfitted reconstructions show up as flat or decreasing cumulative log-posterior and non-monotonic NJSD, catching cases where a gold-standard FSC estimate might still look acceptable.","The method generalizes to any posterior probability of a 3D density given particle images, and the authors point toward refining atomic models against the control set as a next step."],"supporting_citations":[{"why":"Supplies the R-free analogy: an independent test set withheld from fitting is the statistical template the protocol follows.","marker":"[36]"},{"why":"Defines the Bayesian per-image posterior (BioEM) used to score each map against each control particle.","marker":"[37]"},{"why":"Provides the GPU-accelerated BioEM implementation and the integration strategy over orientations and nuisance parameters that the present work optimizes.","marker":"[38]"},{"why":"The Bayesian refinement program used to generate the gold-standard half-set reconstructions whose evolution is monitored.","marker":"[14]"},{"why":"Establishes the overfitting problem and the gold-standard prevention strategy that motivates the control-set test.","marker":"[11]"},{"why":"Sets the gold-standard validation framework (two independent half-set reconstructions) that the protocol is designed to complement.","marker":"[10]"},{"why":"Supplies the RAG1-RAG2 benchmark and the reconstructions against which the pure-noise control set is tested.","marker":"[39]"},{"why":"Supplies one benchmark dataset (HCN1) with known structure and particles used as a non-overfitted test case.","marker":"[40]"},{"why":"Supplies the TRPV1 benchmark dataset used as a non-overfitted test case.","marker":"[41]"},{"why":"Supplies the HIV-1 envelope trimer dataset whose low-resolution reconstruction the authors treat as an overfitted signature case.","marker":"[42]"}],"fun_headline_variants":["Held-out particles expose overfitted cryo-EM reconstructions","Independent particle set catches overfitting in cryo-EM maps","Cryo-EM map validation: one withheld set reveals overfitting","Posterior probability over control set flags bad cryo-EM maps","Control particle set separates genuine cryo-EM maps from noise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The protocol assumes that the BioEM posterior over the control set measures genuine map quality even though the orientations of the control particles are chosen in a first round using the very reconstruction that is being validated, a fit that could let an overfitted map match noise in the control particles and fake the expected trends.","fun_headline_variants_meta":{"raw":{"variants":["Held-out particles expose overfitted cryo-EM reconstructions","Independent particle set catches overfitting in cryo-EM maps","Cryo-EM map validation: one withheld set reveals overfitting","Posterior probability over control set flags bad cryo-EM maps","Control particle set separates genuine cryo-EM maps from noise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1551,"prompt_tokens":993,"completion_tokens":558,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":467}},"tokens_in":609,"tokens_out":558,"duration_ms":5338,"temperature":1.0,"reasoning_tokens":467,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:25:11.737549+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a deliberately overfitted map and rerun the cross-validation with the control particles' orientations fixed by a reference map unrelated to the refinement, or randomized; if the rising cumulative log-posterior and NJSD signature reappears, the discrimination is an artifact of the orientation search rather than a property of genuine maps.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the R-free analogy: an independent test set withheld from fitting is the statistical template the protocol follows."},{"cited_title":"& Hummer, G","cited_arxiv_id":null,"evidence_quote":"Defines the Bayesian per-image posterior (BioEM) used to score each map against each control particle."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the GPU-accelerated BioEM implementation and the integration strategy over orientations and nuisance parameters that the present work optimizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Bayesian refinement program used to generate the gold-standard half-set reconstructions whose evolution is monitored."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the overfitting problem and the gold-standard prevention strategy that motivates the control-set test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sets the gold-standard validation framework (two independent half-set reconstructions) that the protocol is designed to complement."},{"cited_title":"& Others","cited_arxiv_id":null,"evidence_quote":"Supplies the RAG1-RAG2 benchmark and the reconstructions against which the pure-noise control set is tested."},{"cited_title":"& MacKinnon, R","cited_arxiv_id":null,"evidence_quote":"Supplies one benchmark dataset (HCN1) with known structure and particles used as a non-overfitted test case."},{"cited_title":"Structure of the TRPV1 ion channel determined by electron cryo-microscopy","cited_arxiv_id":null,"evidence_quote":"Supplies the TRPV1 benchmark dataset used as a non-overfitted test case."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the HIV-1 envelope trimer dataset whose low-resolution reconstruction the authors treat as an overfitted signature case."}],"review_version":1}