{"id":"7f6aa67f-7b03-4795-8645-3efac602b2ab","arxiv_id":"2502.19431","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Time-assisted MCR with pseudo-cubes resolves sequentially acquired fluorescence EEMs whose concentrations change during acquisition.","lead":"This paper shows that fluorescence data recorded while sample concentrations are changing can still be analyzed with Multivariate Curve Resolution if the readings are grouped into short time windows and the exact measurement times are used to guide smoothing. The method was tested on chromatographic and kinetic fluorescence datasets and recovered spectra and concentration profiles similar to reference values.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MCR success may be inherited from PARAFAC initialization built on the same data; without calibration-only initialization, the reported similarities and REP could reflect leakage.","rationale":"The reader's weakest assumption identified EM imputation bias as the core risk; my concern is closely related but more specific: the initialization itself may already contain information about the validation samples, and the EM imputation is the mechanism that propagates that information. The paper is transparent about needing PARAFAC initialization (Section 3.3.4), but it does not test whether the reported accuracy survives when the initialization is built only from calibration samples. If it does not, the central claim is not that time-assisted MCR independently resolves the data, but that it refines a previously fitted model. The validation samples are external to the calibration model, but not necessarily external to the PARAFAC initialization, so the predictive validation is weaker than it appears. I do not reject the paper because the disclosed limitations and the plausible physical reasoning support a conditional acceptance, pending the calibration-only initialization test and, ideally, release of code/data. The reader's conditional verdict remains appropriate, and my analysis does not move it to a different category.","tokens_in":23800,"tokens_out":7004,"duration_ms":80278,"concrete_test":"Re-run the full pipeline for each validation sample using initial estimates (spectral profiles and imputation) obtained from a 4-way PARAFAC model fit exclusively to the calibration samples; keep all MCR settings identical. If the validation REP and spectral similarities degrade materially (e.g., REP above 10% or similarity below 0.95), the reported results are dominated by leakage from the validation samples into the initialization. As a secondary check, repeat with random or SIMPLISMA initialization to quantify the method's intrinsic convergence behavior.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3.1 states that initial estimates and initial missing-data imputations were taken from 4-way PARAFAC models of the same datasets established in previous studies [5-7]. Section 3.3.4 then reports that without those PARAFAC results (used for both imputation and initialization), MCR faces convergence issues and fails. This creates a leakage path: if the prior PARAFAC models were fit to the full sample set, including the validation samples, then the validation REP (1.76% LC, 3.67% Kin) and spectral similarities above 0.99 are not independent evidence that time-assisted MCR extracts the profiles from the observed data. The EM imputation at 90.9-96.875% missing propagates the initialization into every iteration (Section 2.2), so the final solution may be a refinement of the initialization rather than an independent resolution. The paper's claim that time-assisted MCR resolves such data effectively is therefore conditional on the unstated assumption that the initialization does not already encode the answer. That assumption is not tested.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a Multivariate Curve Resolution (MCR) strategy for sequentially acquired fluorescence Excitation-Emission Matrices (EEMs) recorded while fluorophore concentrations change. The authors introduce 'psCubes', partially filled three-way arrays in which each short time segment of an emission scan is treated as a row with a high fraction of missing entries, and they use Expectation Maximization imputation plus time-based smoothing of concentration profiles. The method is applied to two datasets: LC-EEM (pyridoxine with interferents, chromatographic conditions) and Kin-EEM (diclofenac photodegradation kinetics). For both datasets, the authors report high spectral similarity with reference spectra (>0.99), low validation REP (1.76% and 3.67%), and physically meaningful concentration profiles, despite 90.9% and 96.875% missing data. The central claim is that time-assisted MCR resolves such data effectively.","tokens_in":24144,"tokens_out":8676,"duration_ms":90798,"significance":"If the claim were validated independently, the contribution would be practically useful: it would allow a bilinear MCR framework to handle data that are not conventionally bilinear, leveraging the time information to recover profiles from highly incomplete EEMs. The paper is transparent about several limitations, including the need for good initial estimates, and it reports detailed real experimental data. However, the evaluation does not currently provide independent evidence for the method's ability to resolve the data from the observed measurements alone, and the manuscript overstates the strength of the validation. The approach is potentially valuable as a refinement or downstream analysis tool when prior multilinear solutions are available, but that narrower claim requires reframing and additional experiments.","major_comments":[{"comment":"The central evaluation is not independent of the PARAFAC benchmark. Section 2.3.1 states that the MCR initial estimates and the initial missing-data imputations were taken from 4-way PARAFAC models of the same datasets from previous studies [5-7]. Section 3.3.4 then admits that without these prior results, or when they are used only for imputation or only for initialization, the MCR models 'faced convergence issues' and produced unsatisfactory results. Consequently, the reported REP values (1.76% for LC-EEM, 3.67% for Kin-EEM) and spectral similarities (>0.99) may reflect information inherited from the PARAFAC initialization rather than the ability of time-assisted MCR to resolve the data from the observed measurements. The comparison to 'processing the same data as cubes' is circular because the cube solutions are the MCR starting points. To support the abstract's claim that 'time-assisted MCR resolves such data effectively,' the authors should either (a) fit the prior PARAFAC models to the calibration set only and then apply MCR to the validation samples, (b) demonstrate concretely that MCR improves on the PARAFAC initialization in validation error or profile accuracy, or (c) show that the final MCR solution is insensitive to the choice of initialization. As written, this is a load-bearing gap in the validation.","section":"§2.3.1 and §3.3.4"},{"comment":"The treatment of missing data is insufficiently specified to rule out self-consistent artifacts. With 90.909% (LC-EEM) and 96.875% (Kin-EEM) of the entries missing, the imputed values vastly outnumber the experimental ones in the matrices used by the alternating least squares updates. The paper states that 'the minimization of errors in the fitting function during the iterations only considers the experimental data and never the imputed ones,' but the described procedure (EM imputation followed by bilinear modeling at each iteration, similar to the N-way Toolbox for PARAFAC) normally minimizes the complete-data loss, which includes the imputed entries. If the algorithm instead uses a weighted least-squares objective on the observed entries, that objective and the weighting scheme need to be stated explicitly. Without such clarification, and without a sensitivity analysis with respect to the initial imputation (e.g., random or perturbed initial imputations, or alternative starting profiles), the possibility remains that the resolved profiles are largely determined by the PARAFAC-based initialization rather than by the experimental information. This is particularly concerning because Section 3.3.4 reports convergence failures for other initialization schemes.","section":"§2.2"}],"minor_comments":[{"comment":"The acronym nVarCLFC is used before its full expansion; the expansion 'number of Variables with Constant Local Fluorophore Concentrations' should be given at the first occurrence.","section":"§2.2"},{"comment":"The statement that 'these details have not been shown here and are part of future work' refers to the negative results with alternative initializations; for a journal article, at least a brief summary of the convergence failures and their characteristics should be included in the main text or supplementary information.","section":"§3.3.4"},{"comment":"The abbreviation 'EJCR test OK' in the calibration tables is not defined in the text; please provide a definition or reference.","section":"§3.1 / Table S2, S3"},{"comment":"The argument that the success of the chosen unfolding strategy 'suggests that such interactions do not actually occur' is an inference from a single successful model; a direct test of the EX-EM independence assumption, or at least a comparison with an alternative unfolding, would strengthen this point.","section":"§3.3.2"}],"recommendation":"major_revision","confidential_remarks":"This is a chemometrics paper that fits the journal's applied analytical scope. The main concern is not the authors' transparency but the gap between the abstract's strong claim and the conditional nature of the method as actually demonstrated. If the authors can add a calibration-only initialization experiment or otherwise show that the final solutions are not inherited from the PARAFAC seeds, the paper would be publishable. If such an experiment is not feasible, the claim must be weakened substantially in the abstract and conclusions to describe a refinement strategy that depends on prior multilinear solutions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this as an extension of the authors' own PARAFAC program, not as a standalone method validation. The genuinely new piece is feeding recorded acquisition times into MCR on pseudo-cubes (psCubes) built from sequentially acquired EEMs, with EM imputation each iteration and time-based smoothing of c(t) profiles. The two datasets (LC-EEM and Kin-EEM) are real and the resolved spectra match references well: similarity ≥0.9963/0.9995 for LC, >0.99 for Kin, REP 1.76% and 3.67%. Those are respectable numbers.\n\nThe paper is honest about a big caveat: it needs PARAFAC results from the same data to initialize and to fill the missing entries; without them MCR doesn't converge (Section 3.3.4). That is the soft spot and it is central. With 90.9% and 96.875% missing entries, the EM imputation is driven by the current bilinear model, and the seeds are the very PARAFAC solutions the paper later compares against. The external similarity to reference spectra and the validation REP are independent of that benchmark, so the results are not pure artifacts, but they are not an independent resolution either. The reader should treat the reported quality as conditional on the prior PARAFAC solutions encoding correct spectral shapes.\n\nOther soft spots, minor in comparison: nVarCLFC and smoothing parameters are chosen per dataset on these data; no code or raw data are shipped, so the EM-with-MCR implementation cannot be reproduced externally. The paper acknowledges most of this.\n\nWho is it for: a chemometrician working on chromatographic/kinetic fluorescence with sequential scanning, or anyone using missing-data EM in multiway models. It deserves a serious referee: the idea is sensible, the application is real, and the leak path is clearly disclosed and can be fixed by adding a calibration-only or SIMPLISMA-only initialization comparison. I would send it out, asking for that control.","headline":"Time-assisted MCR on pseudo-cubes is a plausible extension of the authors' PARAFAC work, but the validation inherits too much from the initialization to count as an independent demonstration.","tokens_in":24544,"tokens_out":1563,"would_cite":false,"duration_ms":14581,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fluorescence data can change while an excitation-emission matrix is being recorded and still be resolved into clean spectra and concentration profiles.","keywords":["fluorescence spectroscopy","excitation-emission matrices","multivariate curve resolution","time measurements","missing data imputation","psCubes","chromatography-fluorescence","reaction kinetics"],"falsifier":"Run time-assisted MCR on synthetic excitation-emission data with a known ground-truth time course and randomly delete 90% of the entries, then compare the recovered concentration profile with the true one and with what expectation-maximization would predict from the initial spectra alone; if the recovered profile matches the truth much better than it matches the imputation prior, the imputation is not steering the solution, and if it does not, the reported resolution is partly an artifact.","tokens_in":23600,"feed_emoji":"🔬","tokens_out":5479,"duration_ms":51427,"temperature":0.7,"pith_summary":"This paper tries to show that sequentially recorded fluorescence data remain analyzable even when fluorophore concentrations change during acquisition, by abandoning the idea that each excitation-emission matrix is one bilinear block. Instead, each emission scan is cut into short time fragments, arranged in time into partially filled three-way cubes, and modeled with Multivariate Curve Resolution assisted by measured times. The paper reports that for liquid-chromatography fluorescence data the lowest spectral similarity against reference spectra was 0.9963 for excitation and 0.9995 for emission, with a validation relative error of prediction of 1.76%, while for a kinetics dataset similarity stayed above 0.99 with a 3.67% error, despite 90.9% and 96.875% missing entries respectively. If true, this extends curve resolution to settings where the standard bilinearity assumption fails, and it makes calibration and quantitation possible in flow and reaction systems with ordinary spectrofluorometers.","feed_headline":"Fluorescence signals can shift mid-scan and still be resolved","feed_subtitle":"Time-localized fragments plus curve resolution recover true spectra and profiles even when over 96 percent of the data is missing.","key_machinery":"The central object is the psCube, a three-way array built from pseudo-EEMs: each pseudo-EEM is a matrix of the dimensions of the original EEM that contains only the experimental readings from one short fragment of one emission scan at one excitation wavelength, with everything else marked missing. Working with these cubes lets the model treat a slowly changing concentration as constant inside each fragment while still sampling the whole time course at many points. The argument is carried by three devices used together: time measurements that localize each fragment, expectation-maximization imputation of missing entries from the current bilinear model with the error computed only over experimental data, and smoothness constraints on the concentration profiles evaluated with each sample's measured times.","core_discovery":"The central claim is that the apparent failure of bilinearity in EEMs acquired during chromatography or kinetics is not a property of the data but of the chosen structure. Because each single fluorescence reading is recorded in milliseconds, concentrations can be treated as constant within short intervals; the paper shows that grouping readings into pseudo-EEMs and localizing those fragments with measured times produces partially filled cubes whose reshaped matrices are compatible with bilinear MCR. With non-negativity, unimodality, and time-based smoothing, the resolved excitation and emission spectra match references closely, the concentration profiles are physically meaningful, and calibration predictions are comparable to those from higher-order models applied to the same data.","pith_inferences":["If the psCube idea generalizes, any sequentially acquired spectroscopic data with fast individual reads and a slow drift, such as infrared or Raman reaction monitoring, could be reorganized the same way and resolved by bilinear MCR.","The reported mean recoveries slightly above and below 100%, 101.76% for LC-EEM and 96.33% for Kin-EEM, are consistent with the imputation step gently biasing concentrations toward the model, a point the paper leaves for future work.","A direct stress test would be to run the same workflow on synthetic dynamic EEMs with known true profiles and randomly delete 90% of the entries, which would isolate the effect of imputation from experimental noise."],"forward_implications":["Conventional EEMs are not the only valid unit for fluorescence modeling; the same sequential scans can be read as high-resolution time profiles.","Time-assisted MCR can replace or complement PARAFAC in kinetic and chromatographic systems, giving each sample its own concentration profile instead of a shared shape.","A very high percentage of missing data is not inherently fatal: the structural bilinearity plus time localization carried models with 90.9% and 96.875% missing entries.","Calibration and prediction for an analyte that photodegrades during measurement can be done with the same data used for resolution, without dedicated reaction-monitoring hardware.","Time measurements must be used during the iterations, not only for graphing, because irregular inter-scan gaps are absorbed by time-based smoothing."],"supporting_citations":[{"why":"Supplies the LC-EEM dataset, the prior time-assisted PARAFAC treatment, and the psCube construction logic.","marker":"[5]"},{"why":"Provides the expectation-maximization approach for missing values that the MCR iterations reuse.","marker":"[4]"},{"why":"Supplies the MCR-ALS methodology that the time-assisted implementation modifies.","marker":"[1]"},{"why":"Supplies the MVC3 toolbox base code that was adapted with imputation and time-based smoothing.","marker":"[23]"},{"why":"Supply the Kin-EEM dataset and the reference kinetic profiles used for initialization and comparison.","marker":"[6,7]"},{"why":"Justifies smoothing model profiles based on predictor measurements such as time.","marker":"[22]"},{"why":"Defines PARAFAC and the multilinearity assumptions that motivate the psCube arrangement.","marker":"[2]"},{"why":"Defines the similarity criterion used to score resolved spectra against references.","marker":"[12]"}],"fun_headline_variants":["Time-stamped fragments rescue fluorescence data mid-scan","Rapid chemistry no longer spoils fluorescence matrix analysis","Curve resolution uses time slices to fix shifting signals","Fluorescence scans survive concentration changes with time aids","Partial cubes and time data resolve dynamic fluorescence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole strategy stands on the assumption that imputing the missing entries from the current bilinear model at each iteration does not quietly shape the final answer, because the fitting error is computed only on the roughly 3 to 9 percent of readings that were actually measured.","fun_headline_variants_meta":{"raw":{"variants":["Time-stamped fragments rescue fluorescence data mid-scan","Rapid chemistry no longer spoils fluorescence matrix analysis","Curve resolution uses time slices to fix shifting signals","Fluorescence scans survive concentration changes with time aids","Partial cubes and time data resolve dynamic fluorescence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1443,"prompt_tokens":961,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":409}},"tokens_in":577,"tokens_out":482,"duration_ms":5680,"temperature":1.0,"reasoning_tokens":409,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T18:18:59.409032+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run time-assisted MCR on synthetic excitation-emission data with a known ground-truth time course and randomly delete 90% of the entries, then compare the recovered concentration profile with the true one and with what expectation-maximization would predict from the initial spectra alone; if the recovered profile matches the truth much better than it matches the imputation prior, the imputation is not steering the solution, and if it does not, the reported resolution is partly an artifact.","supporting_citations":[{"cited_title":"Siano, L","cited_arxiv_id":null,"evidence_quote":"Supplies the LC-EEM dataset, the prior time-assisted PARAFAC treatment, and the psCube construction logic."},{"cited_title":"Tomasi, R","cited_arxiv_id":null,"evidence_quote":"Provides the expectation-maximization approach for missing values that the MCR iterations reuse."},{"cited_title":"Jaumot, A","cited_arxiv_id":null,"evidence_quote":"Supplies the MCR-ALS methodology that the time-assisted implementation modifies."},{"cited_title":"Timmerman, H.A.L","cited_arxiv_id":null,"evidence_quote":"Justifies smoothing model profiles based on predictor measurements such as time."},{"cited_title":"Gómez, M","cited_arxiv_id":null,"evidence_quote":"Defines the similarity criterion used to score resolved spectra against references."}],"review_version":1}