REVIEW 3 major objections 5 minor 1 cited by
A pipeline for Megahertz X-ray Photon Correlation Spectroscopy on soft matter samples at the MID instrument of European XFEL
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A standard configuration and a highly automated data-processing pipeline make megahertz X-ray photon correlation spectroscopy at the MID instrument of the European XFEL a near-routine technique for measuring sub-microsecond dynamics in…
desk verdict A useful, honest instrumentation paper describing a real MHz-XPCS pipeline at EuXFEL MID; the main gap is that end-to-end dynamics are never quantitatively validated against a known sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the burst-train acquisition scheme plus the off-correlation correction. Each 100 ms train illuminates a different sample position, so within a train the same spot is probed by up to 352 pulses spaced by hundreds of nanoseconds; the TTCF is computed only for same-train pulse pairs and averaged over trains. The off-correlation between train $t_n$ and $t_{n+1}$ samples a different spot, so it contains no true single-train relaxation signal but does contain the same detector pedestal-jump artifacts; subtracting it from the TTCF removes those artifacts. The second ingredient is computational: after photonization, per-pixel RMS outliers are flagged within q-annuli, data are processed module-by-module with dense or sparse linear algebra depending on measured array density, and caching keeps memory within cluster-node limits.
What would settle it
A static sample with no dynamics should give a flat, zero-baseline corrected correlation function; if any 32x32 block-shaped artifacts survive in the corrected TTCF, the off-correlation subtraction is not removing transient pedestal jumps. Quantitatively, diluted colloidal spheres with a known diffusion constant should yield the same relaxation time before and after correction at average count rates near $10^{-3}$ to $10^{-2}$ photons per pixel per pulse.
Extended reading notes
Core claim
The paper's central claim is that MHz-XPCS can now be conducted and analyzed with minimal user intervention at MID, because a standard configuration and an automated, hardware-aware pipeline replace the manual expert workflows that previously limited the technique. A key step is computing the two-time correlation function per X-ray train, using pulse-pulse pairs within each train while a continuing translation of the sample moves a fresh spot into the beam; this both limits radiation dose and makes each train an independent dynamics measurement. To remove detector artifacts from 'jumping' pixels, the pipeline also computes off-correlations between adjacent trains and subtracts them, which shifts the correlation baseline and largely eliminates square 32x32 memory-cell artifacts in the TTCF. After outlier masking, q-binning, and train-averaging, the final $g_2$ functions are delivered through the DAMNIT interface and prepared for FAIR publication-ready output.
Load-bearing premise
The off-correlation subtraction assumes that the detector's intermittent spurious signals last long enough and appear the same way in pictures taken from neighboring spots, so that subtracting them removes the spurious signal without removing the real motion of the sample.
Editorial extensions
If this is right
- A user at MID can now plan a MHz-XPCS experiment on a soft-matter sample and obtain analyzable train-averaged $g_2$ functions during the beam time, because geometry calibration, outlier detection, intensity filtering, and correlation calculation are automated.
- Radiation-sensitive samples such as protein solutions become practical, since the fresh-spot-per-train scheme spreads dose and the pipeline retains usable correlation data down to an average count of roughly $10^{-3}$ photons per pixel per pulse.
- Residual detector artifacts from pedestal 'jumping' pixels are suppressed by subtracting adjacent-train off-correlations, which also moves the correlation baseline from 1 to 0 and raises the signal-to-noise ratio without spending extra sample or beam time.
- Processed results are exposed through the DAMNIT interface and are being prepared for a FAIR repository, meaning published correlation functions can be traced back to calibration, masks, and train-level metadata.
- The pipeline's benchmark-driven core counts and sparse-matrix mode keep the correlation calculation for multi-hundred-gigabyte runs within the memory of a single cluster node, so no special-purpose supercomputer is required.
Reading between the lines
- Implicit boundary: the pipeline measures only intra-train dynamics, i.e. pulse lags within a 352-pulse window, so slower millisecond-to-second processes are not captured by these same-train TTCFs even though the fresh-spot scheme enables long cumulative measurements.
- A testable check would be to record repeated runs on a stable colloidal standard and compare corrected and uncorrected $g_2$ amplitudes, mapping the range of pedestal-jump durations for which the off-correlation subtraction is safe.
- If event-based detectors with direct hit-list readout (such as the Timepix4 concept the paper cites) become standard, the sparse linear algebra the pipeline already uses could become the default, removing most of the data-volume bottleneck described here.
- Extending this pattern to other burst-mode XFEL facilities would turn sub-microsecond XPCS from a specialist technique into a standard probe of biological and soft-matter dynamics, much as BioSAXS was routinized at third-generation synchrotrons.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents the experimental configuration, data acquisition protocol, and automated data processing pipeline developed for megahertz X-ray Photon Correlation Spectroscopy (MHz-XPCS) at the MID instrument of European XFEL. The pipeline includes mean-intensity calculations, automatic geometry and beam-center refinement, outlier pixel detection and masking, intensity-based filtering of compromised sample positions, computation of two-time correlation functions (TTCFs) and g2 functions using both dense and sparse linear algebra, and an off-correlation subtraction step intended to suppress detector artifacts such as 'jumping' pixels. The authors illustrate the pipeline with ferritin solution data, discuss computational resource management on the DESY Maxwell cluster, describe output tools (DAMNIT, Jupyter notebooks), and outline FAIR-data efforts in the context of DAPHNE4NFDI. The central claim is that MHz-XPCS can now be conducted and analyzed with minimal user intervention and with high data quality suitable for publication.
Significance. If the pipeline performs as claimed, it would make sub-microsecond XPCS on soft matter accessible to non-specialist users at a major XFEL facility, a genuinely useful step for the community. The paper provides a detailed, reproducible description of the data handling workflow, including concrete benchmarks (Figs. 4 and 8), an openly referenced software repository (extra-speckle), and a clear presentation of the correction logic. The explicit treatment of AGIPD-specific artifacts, the sparse-data optimization, and the integration with DAMNIT and FAIR data platforms are valuable practical contributions. However, the paper's central claim of 'high data quality suitable for publication' is not backed by a quantitative end-to-end validation against known sample dynamics; the demonstration is instead based on illustrative images and correlation maps that show visible improvements after each correction. This missing validation is the main factor limiting the paper's impact.
major comments (3)
- [Abstract and Sec. 7] The central claim that the pipeline delivers 'high data quality suitable for publication' with 'minimal user intervention' is not quantitatively validated. The paper never compares a g2(q, Δτ) function produced by the pipeline with the dynamics of a known sample, whether from a fit to a diffusional model with an independently measured diffusion coefficient, from dynamic light scattering, or from a previous XPCS experiment. The ferritin data in Figs. 3, 6, 7, and 12 demonstrate that the correction steps run and improve the visual appearance of the data, but they do not establish that the final correlation functions are quantitatively faithful. Because all the correction steps in Secs. 3.4, 3.5, and 5 could execute successfully and still deliver biased dynamics, a test on a standard sample with known dynamics is load-bearing for the claimed routine usability. I recommend adding an end-to-end validation, including a comparison of extracted decay times (e.g., from g2 - 1 = β exp(-2Dq^2 Δτ)) against literature or an independent measurement, with residuals or confidence intervals.
- [Sec. 5 and Appendix A] The off-correlation subtraction, which is the principal artifact-suppression step, rests on the assumption that detector artifacts such as 'jumping' pixels contribute identically to same-train TTCFs and adjacent-train off-correlations. The Appendix itself acknowledges that jumps lasting only a few trains can evade correction, meaning the subtraction is not guaranteed to remove all artifacts. The magnitude of any residual bias on the extracted dynamics is not quantified anywhere in the paper. Since Sec. 5 states that data with kbar < 10^-3 are excluded because residual artifacts dominate, the reader is left without a quantitative bound on the artifact level in the accepted data range. I recommend quantifying the residual by, for example, injecting simulated jumping-pixel artifacts into a dataset with known dynamics and measuring the resulting bias in g2, or by providing a diagnostic that flags trains/pixels for which the stationarity assumption fails.
- [Secs. 3.4 and 3.5] The outlier detection thresholds (normalized RMS window 0.75–1.75), the intensity filter deviation limit (±20%), and the low-intensity cutoff (kbar < 10^-3) are described as empirically chosen and user-configurable, but no sensitivity analysis is provided. Because these parameters directly decide which pixels and trains contribute to the correlation functions, the 'minimal user intervention' claim depends on the pipeline's results being robust to reasonable variations in these settings. I recommend showing that the final g2 functions are stable when the thresholds are varied over a plausible range, or at least reporting the fraction of pixels/trains rejected in the example dataset so readers can gauge the impact.
minor comments (5)
- [Fig. 5 caption and Sec. 3.3] The word 'intergated' should be 'integrated' in the caption of Fig. 5.
- [Fig. 4 and Fig. 8] The performance benchmarks in Figs. 4 and 8 appear to be single measurements without error bars or repeated-run statistics; please state whether the points are single runs and whether the qualitative conclusions are robust to run-to-run variability.
- [Fig. 12 caption and Sec. 4] There are typographical errors: 'Performace' should be 'Performance' in the Fig. 12 caption, and 'referes' should be 'refers' in Sec. 4.
- [Sec. 6] 'inclduing' should be 'including', and the spaced 'F AIR' should be a single word 'FAIR' in the text.
- [Eq. (3) and Eq. (4)] The notation for the TTCF is inconsistent: Eq. (3) uses (q, τ1, τ2), Eq. (4) writes TTCF(q, τ1, Δτ), and Eq. (5) uses train/pulse indices. Please define the mapping between the continuous-time notation and the discrete train/pulse indices, and clarify the averaging in Eq. (4) over τ1 for a fixed Δτ.
Circularity Check
No circularity found: the pipeline and the off-correlation correction are self-calibration procedures with independently described artifact models, not derivations that reduce to their own inputs.
full rationale
This paper is a facility methods and pipeline description rather than a derivation of a physical result, so the standard circularity patterns do not apply. The only step that could initially resemble a self-referential correction is the off-correlation subtraction in Sec. 5 and Appendix A, where adjacent-train correlations are subtracted from same-train TTCFs to suppress 'jumping' pixel pedestal artifacts. This is not circular: adjacent trains illuminate fresh sample spots, so the off-correlation contains no true dynamics by the experiment's own acquisition design, while the artifact mechanism (32-cell pedestal drifts) is independently characterized in Fig. 11 and the correction effect is demonstrated directly in Fig. 12. The outlier detection in Sec. 3.4 uses mean-intensity RMS statistics relative to q-annulus medians, and the beam-center refinement in Sec. 3.3 maximizes azimuthal-profile overlap; both are internal consistency procedures, not fitted parameters renamed as predictions. The heavy self-citation, including Madsen et al. (2021), Dallari et al. (2021b), and the EuXFEL pipeline documentation, is normal for an instrument paper, and the specific correction is described with equations and figures in the present work rather than being imported solely as an unverified authoritative claim. The skeptical observation that no measured g2 is fitted to a known sample's dynamics is a genuine completeness or validation concern, but it is a missing piece of evidence, not a definitional reduction of a claimed result to its own input. Therefore there is no significant circularity.
Assumptions & free parameters
free parameters (5)
- Outlier RMS thresholds (lower=0.75, upper=1.75) =
0.75 and 1.75 times the median RMS
- Intensity filter deviation threshold =
±20%
- Sparse/dense data density threshold =
5e-2
- Minimum scattering intensity (kbar) =
~1e-3 photons/pixel/pulse
- Number of azimuthal sectors for beam center refinement =
8
assumptions (5)
- domain assumption Each X-ray pulse train illuminates a fresh, undeformed sample spot, with translation within a train below 100 nm.
- domain assumption Detector artifacts ('jumping' pixels) are statistically equal in same-train and adjacent-train correlations.
- domain assumption AGIPD in high-CDS mode provides single-photon sensitivity sufficient for reliable photonization above kbar ~1e-3.
- domain assumption The beam center and momentum transfer are accurately determined by azimuthal symmetry optimization combined with motor-position-based geometry.
- standard math Standard two-time correlation function formalism applies to the burst-mode data.
Cite this review
Pith. "Pith review of A pipeline for Megahertz X-ray Photon Correlation Spectroscopy on soft matter samples at the MID instrument of European XFEL." pith.science (2026). https://pith.science/paper/ZRSCCUGW
@misc{pith2026250608668,
author = {Pith},
title = {Pith review of: A pipeline for Megahertz X-ray Photon Correlation Spectroscopy on soft matter samples at the MID instrument of European XFEL},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZRSCCUGW}},
note = {Machine review of arXiv:2506.08668}
}
read the original abstract
In this paper we present the experimental protocol and data processing framework for Megahertz X-ray Photon Correlation Spectroscopy (MHz-XPCS) experiments on soft matter samples, implemented at the Materials Imaging and Dynamics (MID) instrument of the European X-ray Free-Electron Laser (EuXFEL). Due to the introduction of a standard configuration and the implementation of a highly automated data processing pipeline, MHz-XPCS measurements can now be conducted and analyzed with minimal user intervention. A key challenge lies in managing the extremely large data volumes generated by the Adaptive Gain Integrating Pixel Detector (AGIPD) - often reaching several petabytes within a single experiment. We describe the technical implementation, discuss the hardware requirements related to effective parallel data processing, and propose strategies to enhance data quality, in particular related to data reduction strategies and an improvement of the signal-to-noise ratio. Finally, we address strategies for making the processed data FAIR (Findable, Accessible, Interoperable, Reusable), in alignment with the goals of the DAPHNE4NFDI project.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
Depletion-Induced Interactions Modulate Nanoscale Protein Diffusion in Polymeric Crowder Solutions
Ferritin diffusion in polymer crowder solutions follows a c*-normalized non-monotonic curve with a crossover near 2c*, attributed to depletion-induced intermediate-range order that bulk viscosity cannot explain.
Reference graph
Works this paper leans on
-
[1]
Acerbo, A. S., Cook, M. J. & Gillilan, R. E. (2015).Journal of synchrotron radiation,22(1), 180–186. Allahgholi, A., Becker, J., Delfs, A., Dinapoli, R., Goettlicher, P., Greiffenberg, D., Henrich, B., Hirsemann, H., Kuhn, M., Klanner, R.et al.(2019a).Synchrotron Radiation,26(1), 74–82. Allahgholi, A., Becker, J., Delfs, A., Dinapoli, R., G¨ ottlicher, P....
arXiv 2015
-
[12]
IUCr macros version 2.1.10: 2016/01/28
Performace of the off-correlation correction for a selected q-bin, containing the ”jumping” pixel shown in Fig.11: (a) train-averaged TTCF without correction; (b) train-averaged off-correlation; (c) train-averaged TTCF after subtraction of the off-correlation (with this step the baseline shifts to 0). IUCr macros version 2.1.10: 2016/01/28
work page 2016
-
[287]
Origins of suppressed self-diffusion of nanoscale constituents of a complex liquid
Schmidt, P., Ahmed, K., Danilevski, C., Hammer, D., Rosca, R., Kluyver, T., Michelat, T., Sobolev, E., Gelisio, L., Maia, L.et al.(2024).Frontiers in Physics,11, 1321524. Sobolev, E., Schmidt, P., Malka, J., Hammer, D., Boukhelef, D., M¨ oller, J., Ahmed, K., Bean, R., Berm´ udez Mac ´ ıas, I. J., Bielecki, J.et al.(2024).Frontiers in Physics,12, 1331329....
work page Pith review arXiv 2024
-
[646]
Fujisawa, T., Inoue, K., Oka, T., Iwamoto, H., Uruga, T., Kumasaka, T., Inoko, Y., Yagi, N., Yamamoto, M. & Ueki, T. (2000).Journal of applied crystallography,33(3), 797–800. IUCr macros version 2.1.10: 2016/01/28 31 Girelli, A., Bin, M., Filianina, M., Dargasz, M., Anthuparambil, N. D., M¨ oller, J., Zozulya, A., Andronis, I., Timmermann, S., Berkowicz, ...
arXiv 2000
-
[1710]
D., Girelli, A., Begam, N., Kowalski, M., Retzbach, S., Senft, M
Timmermann, S., Anthuparambil, N. D., Girelli, A., Begam, N., Kowalski, M., Retzbach, S., Senft, M. D., Akhundzadeh, M. S., Poggemann, H.-F., Moron, M., Hiremath, A., Gutm¨ uller, D., Dargasz, M., Oeztuerk, O., Paulus, M., Westermeier, F., Sprung, M., Ragulskaya, A., Zhang, F., Schreiber, F. & Gutt, C. (2023).Scientific Reports,13(1), 11048. URL:https://d...
-
[3962]
(1990).Quantum Optics: Journal of the European Optical Society Part B,2(4),
Schatzel, K. (1990).Quantum Optics: Journal of the European Optical Society Part B,2(4),
work page 1990
- [5528]
-
[5580]
Ashiotis, G., Deschildre, A., Nawaz, Z., Wright, J. P., Karkoulis, D., Picca, F. E. & Kieffer, J. (2015).Journal of applied crystallography,48(2), 510–519. Barty, A., Gutt, C., Lohstroh, W., Murphy, B., Schneidewind, A., Grunwaldt, J.-D., Schreiber, F., Busch, S., Unruh, T., Bussmann, M., Fangohr, H., G¨ orzig, H., Houben, A., Kluge, T., Manke, I., L¨ utz...
Show all 10 references
-
[6179]
Lehmk¨ uhler, F., Dallari, F., Jain, A., Sikorski, M., M¨ oller, J., Frenzel, L., Lokteva, I., Mills, G., Walther, M., Sinn, H., Schulz, F., Dartsch, M., Markmann, V., Bean, R., Kim, Y., Vagovic, P., Madsen, A., Mancuso, A. P. & Gr¨ ubel, G. (2020).Proceedings of the National ...
2020 doi
-
[8037]
DESY, (2025a)
Decking, W., Abeghyan, S., Abramian, P., Abramsky, A., Aguirre, A., Albrecht, C., Alou, P., Altarelli, M., Altmann, P., Amyan, K.et al.(2020).Nature photonics,14(6), 391–397. DESY, (2025a). Documentation for the maxwell hpc cluster. URL:https://docs.desy.de/maxwell/ DESY, (202...
2020
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.