Pith. sign in

REVIEW 3 major objections 6 minor 2 references

Solving Crystal Structures by Carrying Out the Calculation of the Single-Atom R1 Method in a Lottery Mode

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A random split of the current atomic model into two child models lets the single-atom R1 method solve crystal structures without a human choosing the next partial model.

desk verdict A genuine attempt at making the sR1 method push-button, but the lottery's acceptance gate is the same approximate R1 it is minimizing, and the paper never shows that lower sR1 minima track structural correctness. read the letter →

arxiv 2412.18625 v2 pith:YQ5KHAZI submitted 2024-12-19 physics.chem-ph physics.comp-phphysics.data-an

classification physics.chem-phphysics.comp-phphysics.data-an
keywords single-atomR1lotteryschemecrystalstructuresolutionstatisticalfluctuationapproximateX-raycrystallographyrandomsplitting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a crystal structure can be solved by running the single-atom R1 (sR1) calculation without any human decisions between cycles, if each cycle's finished model is randomly split into two child models and only the child whose expansion lowers the minimum approximate R1 is kept. The sR1 method locates atoms one at a time using an approximate R1 residual; previously a user had to choose recognizable fragments or delete ghost atoms to start the next cycle. The lottery replaces that judgment with random splits, betting that statistical fluctuation will sometimes produce a small child rich in correctly placed atoms. On test structures with 128 to 316 non-hydrogen atoms per cell, the scheme drove calculations to or close to the correct solutions, and about 85 datasets were tested in lottery mode. The paper also reports speed-ups from sharpened intensities and coarser coordinate precision, and a local search over starting positions for difficult cases.

What carries the argument

The carrying object is the single-atom R1 (sR1), an approximate version of the traditional crystallographic R1 in which only one undetermined atom's coordinates are optimized while terms containing the other undetermined atoms' coordinates are deleted, though their scattering factors are partly retained. The new mechanism is the lottery split: a parent model is randomly partitioned into a deliberately small child (1 to 20 atoms) and a large child, and the minimum approximate R1 reached during expansion is the objective and acceptance criterion. This criterion is what lets the calculation decide whether a child is an improvement without human inspection. Two implementation changes support the scheme: sharpened intensities replace raw intensities, and single-atom positions are refined to 0.2 Å instead of 0.001 Å.

What would settle it

On a set of datasets with known solutions, take a partially correct model, record the minimum approximate R1 reached after one cycle, then perturb the model by displacing atoms by 0.2, 0.5, 1.0, and 2.0 Å; if the minimum approximate R1 does not increase with displacement, the acceptance criterion can reward models that are farther from the truth.

Watch

Extended reading notes

Core claim

The central claim is that the lottery mode can drive an sR1 calculation toward correct structure solutions automatically. After one expansion cycle, atoms whose addition raised the approximate R1 are deleted; the remaining parent model is randomly split into a small child (about 1 to 20 atoms) and a large child. The small child is deliberately small so that random fluctuation can make it nearly all good atoms (producing a rebuild) or nearly all bad atoms (leaving its large sibling slightly richer in good atoms than the parent). The child model whose expansion reaches a lower minimum approximate R1 than the parent becomes the next parent; if neither improves, the unchanged parent continues. The paper reports that in detailed tests the final models matched reference solutions to within 0.5 Å for essentially all atoms, and that repeated runs take different random paths but reach approximately the same final model.

Load-bearing premise

The whole method rests on trusting that every time the approximate R1 number goes down, the model really is closer to the true atomic arrangement.

Editorial extensions

If this is right

  • An sR1 calculation can in principle run unattended, removing the need for a crystallographer to recognize fragments or delete ghost atoms between cycles.
  • Difficult cases that previously required an oriented known fragment can be started from a single atom, at the cost of many more lottery cycles; the paper reports one case needing 243 cycles versus 36 with a fragment start.
  • Low data resolution makes structures harder but not impossible: a structure that needed one cycle at 0.77 Å resolution was solved after 113 lottery cycles when truncated to 1 Å.
  • The method is highly parallel because the sR1 map can be evaluated at all grid points simultaneously, so fast completion depends mainly on available computing power.
  • After the primary structure is found, extra lottery cycles can improve the model, and a bond-length-guided sR1 variant can fill in missing atoms when ghost sites compete.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The acceptance rule assumes the minimum approximate R1 tracks closeness to the true structure; testing that monotonicity directly on benchmark structures would either support or undermine the lottery's foundation.
  • The small-child/large-child split resembles an explore-versus-refine balance: the small child can restart from near scratch, while the large child makes incremental corrections; combining this with automated fragment recognition could make the intelligent and lottery modes complementary.
  • The paper's local search for an optimal starting atom position suggests a universal initialization strategy: run one sR1 cycle from each of 64 nearby grid points and keep the lowest minimum approximate R1, an idea the author states is still at an initial stage.
  • The paper explicitly acknowledges that sR1 cannot distinguish a solution from its inverted image and that success is judged by chemical recognizability; an unattended pipeline would need to resolve that ambiguity before calling a structure solved.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a "lottery mode" for the single-atom R1 (sR1) crystal structure solution method. In each cycle, a current partial model is randomly split into a small child (1-20 atoms) and a large child; each child is expanded by sR1 cycles, and the model with the lower minimum approximate R1 is carried forward. The stated goal is to remove the need for a user to select fragments or delete ghost atoms between cycles, making the calculation "care-free." The author reports that the lottery mode solved four benchmark structures (samples 1-4) with final models compared quantitatively to SHELXT-derived references, and that about 85 datasets were tested in lottery mode with qualitative validation. The central claim is that the lottery scheme can drive an sR1 calculation toward a correct structure solution.

Significance. If the central claim were fully established, this would be a notable practical advance: it would automate the sR1 method's cycle-to-cycle model selection, which previously required human judgment, and it introduces a stochastic splitting idea (small-child fluctuations) that is original in this context. The paper also provides source code on GitHub, which is a genuine strength for reproducibility, and the quantitative comparison of final models against SHELXT results for four samples is a useful check. However, the current evidence does not yet establish the claimed mechanism, because the acceptance criterion is the same approximate R1 being minimized and there is no independent validation of the intermediate decisions. The significance is therefore conditional on additional evidence linking sR1 minima to structural correctness.

major comments (3)
  1. [Section 4 (with Section 2)] The acceptance gate for retaining a child model is the minimum approximate R1 reached during expansion, and Section 4 states: "the minimum approximate R1 ever reached is a good measure of how good the resulting model is." This is the same quantity that the sR1 search minimizes. Section 2, however, concedes that the sR1 is an ad hoc approximation and is "not even implicitly related to an electron density of some partial structure." The paper provides no evidence that lower sR1 minima track structural correctness; the external SHELXT comparison is applied only to the final models of selected runs, often after manual intervention, not to the intermediate models on which each lottery decision is made. Consequently, the observed drops in sR1 during lottery cycles are equally consistent with greedy descent into a wrong-model minimum, and the data do not demonstrate that the lottery preferentially propagates models with more "good" atoms. This is load-bearing for conclusion (1) of Section 6.7 and needs to be addressed, for example by comparing intermediate models at successive lottery cycles against the SHELXT reference.
  2. [Sections 5.3 and 5.4] The four quantified examples do not support the abstract's "care-free" claim. For sample 3, the text reports that the lottery-cycle result contained ghost atoms that were manually deleted, that atom types were manually corrected, and that missing atoms were found by a separate bond-length-guided sR1 step (Section 5.3, Figure 3). For sample 4, ghost C atoms were deleted after step 2 and the missing C atoms were again found by the bond-length-guided step (Section 5.4, Figure 4). Thus, for the two difficult cases, the lottery mode alone did not produce the final correct model; manual intervention and auxiliary methods were required. The conclusion in Section 6.7 should be restricted to what is demonstrated: the lottery mode contributed to solving these structures within a workflow that still contains user intervention.
  3. [Section 6.4] The successful single-atom start for sample 3 relies on a post hoc choice of the initial atom position. The default position (0.3, 0.3, 0.3) failed, and (0.325, 0.3042, 0.3) was selected after a 64-point local grid search motivated by that failure. This is a selected trial, not a pre-specified procedure, and it does not provide evidence for a universal care-free starting method. The paper itself says the idea "currently is only at the initial stage," but this qualification is not carried into the general conclusion of Section 6.7, which states that the lottery mode can drive sR1 toward correct solutions without mentioning the dependence on this tuned starting position.
minor comments (6)
  1. [Section 3] The choice of 0.2 Å coordinate precision alongside a 0.4 Å grid step is not justified; the relationship between the grid spacing and the refinement precision could be clarified.
  2. [Section 4] The splitting algorithm uses N both for the parent model size and for the total number of atoms in the unit cell (Section 2), which could confuse readers; consider renaming one of them.
  3. [Section 5.1] The quantitative comparison with the correct model uses a 0.5 Å criterion, but this is defined only in Section S1; the main text should state the tolerance and the fact that H atoms are excluded.
  4. [Section 6.2] The claim that small child models "provided most of the improvement" is based on only four samples and is an observed tendency, not a general result; a more cautious phrasing would be appropriate.
  5. [Section 6.6 and S9] For the roughly 85 datasets tested in lottery mode, success is assessed by qualitative visual inspection only, and the table in S9 does not indicate which of these required manual intervention or bond-length-guided steps; adding this information would make the scope of the lottery-mode claim clearer.
  6. [Throughout] The paper repeatedly uses "care-free" to describe the method; a more precise term such as "unsupervised" or "automatic" would better match the actual workflow, which still includes manual steps in the difficult cases.

Circularity Check

1 steps flagged · score 2.0 of 10

Internal acceptance gate and 'success' metric are the same approximate R1 being minimized, making intermediate improvement reports definitional; final external SHELXT comparisons keep the central claim partly independent.

  1. self definitional [Section 4 (acceptance criterion) and Section 6.1 (success rate)]
    "By intuition, the minimum approximate R1 ever reached is a good measure of how good the resulting model (after excluding the likely bad atoms as described above) is. This standard has been adopted for comparing models. [...] Among those 33 cycles, 26 led to a drop of the approximate R1 (shown on the charts), while 7 were skipped because of failing to improve the model. So, the success rate of the lottery cycles in this case was 26 over 33."

    'Improvement over model 0' is defined as a lower minimum approximate R1 (Section 4), and a 'successful' lottery cycle is counted as one that 'led to a drop of the approximate R1' (Section 6.1). The minimum sR1 is also the quantity that each expansion minimizes, so accepting a child and measuring success use the same function by construction. Accordingly, Section 6.2's statement that small child models provided 'most of the improvement on the solution, namely providing the most drop in the approximate R1' restates the selection rule rather than independently showing progress toward the true structure.

full rationale

The paper's central derivation is not circular in the strongest sense: the sR1 formula is inherited from the author's prior published work by ordinary citation, and the final validation of the lottery mode for samples 1-4 compares the resulting models against 'correct' models produced independently by SHELXT/SHELXL (Section 5). The acceptance criterion in Section 4, however, equates 'improvement' with a drop in the minimum approximate R1, and Section 6.1 counts the same drops as 'successes'; those internal statistics are therefore tautological with respect to the search objective. Because the paper does provide external final-model checks, this definitional shortcut does not by itself force the central conclusion, but it does make the intermediate 'success rates' non-evidential for structural correctness. The exploratory optimal-start search (Section 6.4) is in-sample and acknowledged as preliminary, so it is not treated as a load-bearing circular step. Overall score 2.

Assumptions & free parameters 7 free parameters · 6 assumptions · 0 invented entities

The central claim rests mainly on heuristic choices (child size, grid step, precision, acceptance criterion) and on the inherited sR1 approximation. There are no new physical entities. The post hoc optimal starting position for sample 3 is the clearest fitted parameter.

free parameters (7)
  • small child size cap (p1) = p1 = min(20/N, 0.5)
    Controls the lottery split: one child is designed to be 1 to 20 atoms so statistical fluctuations can dominate (Section 4).
  • grid step size s = 0.4 Å
    Spacing of the fixed grid on which sR1 is evaluated; influences whether the sR1 holes are caught by grid points (Section 6.4).
  • coordinate precision for single atoms = 0.2 Å
    Replaces the previous 0.001 Å precision to cut calculation time (Section 3).
  • Wilson B factor = per dataset
    Used to sharpen intensities before sR1; estimated from data by the Wilson method (Section 3).
  • C-C bond length for bond-length guided sR1 = 1.39 Å plus 0.3 Å shell
    Guides placement of missing atoms in samples 3 and 4 (Section 5.3).
  • optimal starting atom position for sample 3 = (0.325, 0.3042, 0.3)
    Found by testing 64 local grid points around the default (0.3, 0.3, 0.3) after sample 3 was identified as difficult; post hoc tuning (Section 6.4).
  • model comparison cutoff = 0.5 Å
    Used to count correctly located atoms in validation; chosen by hand (Section 5 and S1).
assumptions (6)
  • domain assumption The single-atom R1 approximation defined in Zhang and Donahue (2024) is a valid target for locating atoms.
    The whole calculation reuses the prior sR1 formula; this paper recaps the idea but does not re-derive or independently benchmark the approximation.
  • domain assumption An atom that causes sR1 to rise when added is likely incorrectly positioned.
    Section 4 states this was verified on about 224 datasets, but the verification is not shown.
  • domain assumption The minimum sR1 reached is a good measure of model quality for comparing models.
    Section 4 adopts this as the improvement criterion; no independent correlation with structural correctness is given.
  • domain assumption Randomly splitting a parent model produces children whose fraction of good atoms fluctuates enough to improve the search.
    Section 4 gives the statistical argument; it assumes good and bad atoms are distributed randomly among atoms in the model.
  • domain assumption SHELXT plus SHELXL refinement provides a correct reference model for validation.
    Section 5 defines the correct model via SHELXT/SHELXL in P1; this is a reasonable external benchmark but is assumed correct.
  • domain assumption Visual recognition of a chemically sensible model is valid evidence that a structure is solved.
    Section S11 defends visual inspection against a reviewer's objection; it is the primary success criterion for the 85 lottery datasets.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Solving Crystal Structures by Carrying Out the Calculation of the Single-Atom R1 Method in a Lottery Mode." pith.science (2026). https://pith.science/paper/YQ5KHAZI

@misc{pith2026241218625,
  author       = {Pith},
  title        = {Pith review of: Solving Crystal Structures by Carrying Out the Calculation of the Single-Atom R1 Method in a Lottery Mode},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YQ5KHAZI}},
  note         = {Machine review of arXiv:2412.18625}
}
read the original abstract

As originally designed [Zhang & Donahue (2024), Acta Cryst. A80, 2370248.], after one cycle of calculation, the single-atom R1 (sR1) method required a user to intelligently determine a partial structure to start the next cycle. In this paper, a lottery scheme has been designed to randomly split a parent model into two child models. This allows the calculation to be carried out in care-free manner. By chance, one child model may have higher amounts of "good" atoms than the parent model. Thus, its expansion in the next cycle favors an improved model. These "lucky" results are carried onto the next cycles. while "unlucky" results in which no improvements occur are discarded. Furthermore, unchanged models are carried onto the next cycles in those "unlucky" occasions. On average a child model has the same fraction of "good" atoms as the parent. Only a substantial statistical fluctuation results in appreciable deviation. This lottery scheme works because such fluctuations do happen. Indeed, test applications with the computing power accessibly by the author have demonstrated that the designed scheme can drive an sR1 calculation to or close to reaching a correct structure solution.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

2 extracted references · 2 canonical work pages

  1. [1]

    Spek, A.L. (2020). Acta Cryst. E76, 1-11

  2. [2]

    & Donahue, J

    Zhang, X. & Donahue, J. P. (2024). Acta Cryst. A80, 237-248

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.