{"id":"6b63be27-1a8f-48ea-a6d2-090bbde2bef6","arxiv_id":"2507.15772","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"DIVA uses a variational autoencoder on first-derivative Raman spectra to cluster plant stress states and identify significant peaks without manual preprocessing.","lead":"The authors introduce DIVA, a variational autoencoder pipeline that processes raw Raman spectra from plant leaves without fluorescence removal or manual peak picking, and they apply it to several abiotic and biotic stresses across three plant species. The potential payoff is a more automated, less biased route to early plant stress detection in agriculture.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'quantitative' peak areas are not Raman peak areas: A(v~) is computed as summed absolute values of the derivative D between zero crossings, i.e., total variation, not the integral of the original peak. Without calibration, the concentration claim is unsupported.","rationale":"The reader's REJECT verdict is appropriate. My concern is more specific than the reader's weakest assumption about fluorescence suppression: it targets the definition of the quantity used for quantification. The Methods text is explicit that A(v~) is a sum of absolute derivative values between zero crossings, which is total variation, not the area under the Raman band. This is not a matter of disagreement with consensus; it is an internal inconsistency between the claim and the operational definition. If the authors intended A(v~) to represent peak area, the manuscript needs a corrected derivation and calibration. If they intended a nonstandard proxy, then the claim of unbiased quantification is unsupported without validation against known concentrations. I note the paper's strengths: the architecture is described in sufficient detail to reproduce, the datasets span multiple species and stressors, and the zero-crossing detection procedure is deterministic. However, all main results are reported on the training set, no held-out test results or error bars appear, and the biological interpretations rely on previously assigned Raman bands. Given the central claim is about a fully automated, unbiased, quantitative workflow, the uncalibrated and mis-specified area metric is the most load-bearing defect. My verdict therefore remains REJECT, with no change to the reader's recommendation.","tokens_in":24725,"tokens_out":4053,"duration_ms":51871,"concrete_test":"Acquire a dilution series of a known Raman-active standard (e.g., β-carotene) in a strongly fluorescent plant-like matrix using the same 830 nm benchtop system, run the full DIVA pipeline on each concentration, and extract A(v~) at the standard's main peak. If A(v~) does not scale linearly (or at least monotonically) with known concentration across at least a 10-fold range, the claimed quantitative measure fails. As a complementary check, recompute the reported A(v~) values in Fig. 2b as true integrated peak areas after standard baseline correction; if the rankings or relative magnitudes change materially, the reported trends are artifacts of the derivative-total-variation definition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that DIVA \"quantifies\" biomolecular concentrations rests on the quantity A(v~), described in Results as \"the area under the curve\" at Raman peak positions. The Methods section, however, defines A differently: for each positive-to-negative zero-crossing of the derivative spectrum D(v~)=dI/dv (Eq. 1), \"the area was computed as the sum of the absolute signal values between the immediately preceding and succeeding zero-crossing points.\" That sum is the total variation of the reconstructed derivative across the peak region, approximately equal to the sum of the two peak-height excursions to adjacent minima, scaled by the discrete wavenumber step. It is not the integrated area under the original Raman band, which is the quantity normally proportional to concentration. Because the peaks are also detected on VAE-reconstructed cluster-median derivatives rather than on measured spectra, A(v~) additionally inherits any reconstruction bias or smoothing. Even if the first-derivative assumption about fluorescence suppression holds perfectly, this metric does not establish quantitative biomolecular concentration. The paper provides no calibration curve, no comparison with conventional baseline-corrected peak areas, and no error bars or statistical tests on A(v~). This is a load-bearing failure: the headline contribution is an automated unbiased quantification, and the definition of the reported quantity does not support that claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces DIVA, a variational-autoencoder-based workflow for analyzing raw Raman spectra of plant tissue without manual baseline correction or a priori peak selection. The method computes the first derivative of the raw spectrum, trains a VAE to embed the derivative spectra in a two-dimensional latent space, reconstructs characteristic derivative spectra from cluster medians, and identifies 'significant' peaks via positive-to-negative zero crossings of the reconstructed derivative. The area A(ṽ) around each zero crossing is claimed to provide a quantitative measure of biomolecular concentration. The authors apply DIVA to light stress, shade avoidance, high-temperature stress, and bacterial infection in Arabidopsis, Choy Sum, and Kai Lan, and report consistent sets of peaks assigned to carotenoids, cellulose/lignin/protein, and pectin.","tokens_in":25044,"tokens_out":7339,"duration_ms":75368,"significance":"If the central claims were fully supported, this would be a useful contribution: an unsupervised, interpretable pipeline that removes the need for manual baseline correction and peak selection in Raman-based plant phenotyping, with demonstrations across multiple species and stressors. The workflow is simple and the latent-space visualizations are intuitive. However, the paper's headline claim of unbiased quantitative analysis is not yet supported: the area metric is not a calibrated concentration measure, the main results are reported only on the training set, and no statistical uncertainty or significance testing is provided. The authors also credit the ability to identify biologically meaningful peaks, but the validation relies on a reference with overlapping authorship. With additional analysis (test-set evaluation, recalibration or redefinition of the area metric, error bars and significance tests, and independent validation), the method could become a valuable tool.","major_comments":[{"comment":"The quantity A(ṽ) described in the Results as 'the area under the curve' and as providing 'a quantitative measure of the concentrations of the biomolecules' is defined in Methods as the sum of the absolute signal values of the derivative D(ṽ) between adjacent zero crossings. This is the total variation of the reconstructed spectrum over the peak region (approximately twice the peak height for a symmetric peak), not the integrated area under the Raman band, which is the quantity proportional to concentration. The manuscript provides no calibration curve, no comparison with conventional baseline-corrected peak areas, and no error bars on A(ṽ). As written, the quantitative-concentration claim is unsupported. Please recompute peak areas from the de-transformed spectra or explicitly reframe A(ṽ) as a peak-height proxy and revise the text and abstract accordingly.","section":"Results: 'DIVA to analyze Raman spectra'; Methods: 'Significant Peaks Detection Methodology'"},{"comment":"The manuscript states that 'results presented in the main figures are based on these training sets' and that the remaining 10% was 'held out for testing and validation,' yet no test-set results are shown anywhere. For a VAE, training-set embeddings can reflect reconstruction of the training data rather than generalizable clustering; without projecting the held-out spectra into the trained latent space and showing that the cluster medians and identified peaks are stable, the generalizability and 'unbiased' claims are not established. Please include test-set latent projections, cluster assignments, and reconstructed spectra for at least the main case studies.","section":"Methods: 'Training set sizes and Model execution times'"},{"comment":"All values of A(ṽ) are presented as point estimates without error bars, confidence intervals, or significance tests. Because multiple leaves, locations, and replicate spectra were acquired per condition, the authors should report the distribution of A(ṽ) (or the corrected area metric) across biological replicates and test whether differences between conditions are statistically significant. The fixed choice of the 'five most significant peaks' is an arbitrary cutoff and should be replaced by a statistical threshold or justified explicitly.","section":"Figs. 2–5 and Extended Data Fig. 1"},{"comment":"The biological validation of the identified peaks relies on reference [15], which shares authors with the present study and which reports the same Raman peak assignments (carotenoids at 1150–1521 cm⁻¹, cellulose/lignin/protein at ~1318 cm⁻¹, pectin at ~743 cm⁻¹). Because the same peak set is 'discovered' in every stress condition and is then validated against a non-independent reference, the claim that DIVA identifies previously unknown or unbiased markers is not supportable. Please provide independent validation, such as comparison with a conventional baseline-corrected analysis or published peak assignments from a different group, and state how many of the top-five peaks differ from the reference set.","section":"Results: 'Decoding light stress responses…'; reference [15]"},{"comment":"The abstract and introduction claim that DIVA processes 'native Raman spectra—including fluorescence backgrounds—without manual preprocessing.' In practice, the pipeline includes normalization ('The spectra were normalized by reducing the original Raman intensities'), cosmic-ray removal, and spectral trimming, so the claim should be qualified. More substantively, the core assumption that the first derivative suppresses the fluorescence background 'without requiring ad hoc background fitting procedures' is not tested on spectra with strong or steeply varying fluorescence; if the fluorescence gradient is large relative to the Raman peaks, the derivative input will be dominated by the background. Please include a quantitative characterization of the fluorescence background in the datasets or a simulated stress test with varying fluorescence slopes.","section":"Data preprocessing; Results: first-derivative fluorescence suppression"}],"minor_comments":[{"comment":"The description 'Leaky ReLU activation layer with an alpha value of 1' corresponds to the identity function, not a leaky nonlinearity; please clarify whether the decoder actually uses a nonlinear activation.","section":"Methods: 'Variational Autoencoder for Unsupervised Clustering'"},{"comment":"The encoder and decoder descriptions do not specify the latent dimension explicitly; the captions refer to a 'two-dimensional latent space,' but the layer with 4 neurons preceding the sampling layer suggests a latent dimension of 2. Please state the latent dimension explicitly.","section":"Methods: 'Variational Autoencoder for Unsupervised Clustering'"},{"comment":"The supplementary figures refer to 'detransforming the derivative D(ṽ) spectra'; please explain the integration procedure and the constant of integration used to recover I(ṽ).","section":"Supplementary figures"},{"comment":"The paper does not include a data or code availability statement; given the emphasis on reproducibility and unbiased analysis, please provide access to the processed spectra and the implementation.","section":"General"},{"comment":"Several biological interpretations (e.g., the 'oscillating immune behavior' in Extended Data Fig. 1) are speculative; consider phrasing these as hypotheses rather than conclusions.","section":"Discussion"},{"comment":"Reference [28] is a preprint; if a peer-reviewed version exists, please cite it.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The reliance on reference [15] for validation, the absence of data/code availability, and the mismatch between the paper's framing (cs.LG method paper) and the biological claims may require editorial attention. If the authors can recalibrate the area metric and provide test-set and uncertainty analyses, the manuscript could become suitable for publication; otherwise, the quantitative claims should be substantially weakened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is reasonable: take first derivatives of raw Raman spectra to suppress fluorescence baselines, train a VAE on those derivatives, and use zero-crossings in reconstructed cluster medians to locate stress-responsive peaks without manual selection. Applying this across three species and several stressors is a legitimate new application, and the latent-space separations look sensible. The paper is clearly written and the workflow is reproducible enough that I could re-implement it.\n\nThe soft spots are real and fairly serious. The quantity called the area under the curve, A(ν̃), is not the area under the original Raman peak. Per Methods, it is the sum of absolute derivative values between adjacent zero-crossings—i.e., total variation of the derivative, which scales with peak height/width, not height×width. That means the claim that DIVA \"quantifies\" biomolecular concentrations is unsupported, and without a calibration curve or comparison to conventional baseline-corrected areas, the concentration language should be revised or dropped.\n\nThe evidence is also thinner than the text implies. All main figures come from the training set; the 10% held-out test set is mentioned but never shown. There are no error bars or significance tests on A(ν̃). The \"top five\" peaks cutoff is hand-picked, and the biological assignments rely on reference [15], which shares authors with this paper—not fatal, but worth flagging. The fluorescence-suppression assumption is stated but not stress-tested on spectra with severe or structured backgrounds.\n\nThat said, the paper is not incoherent. The identification and clustering side could still be useful even if the quantification is wrong, and the flaws are the kind a referee could push the authors to fix: recalibrate the metric, show test-set performance, add error bars, and benchmark against standard preprocessing. The data are real and the experiments are careful.\n\nFor a reading group, it is a decent cautionary example about metric definitions in spectral analysis. I would not cite it as a quantitative method in its current form. But I would send it to peer review rather than desk-reject: the core pipeline is worth evaluating seriously, and a competent referee can drive the revisions needed to make the claims match the method.","headline":"A plausible VAE-on-derivative pipeline for plant stress Raman spectra, but the 'quantitative' peak areas are total variation of the derivative, not Raman peak areas, and the test set never appears.","tokens_in":25528,"tokens_out":2244,"would_cite":false,"duration_ms":27519,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A variational autoencoder fed only with the first derivative of raw Raman spectra can detect and quantify plant stress across species and stressors without baseline correction, normalization, or pre-chosen peaks.","keywords":["Raman spectroscopy","variational autoencoder","plant stress detection","unsupervised spectral analysis","fluorescence background suppression","deep learning","precision agriculture","zero-crossing peak detection"],"falsifier":"Simulate or collect Raman spectra with known peaks embedded in a steep, structured fluorescence background, such as an exponential or sharply sloped baseline, run them through the DIVA workflow, and check whether the detected peak positions and relative areas match the known inputs; spurious zero-crossing peaks, shifted positions, or noise-dominated reconstructions would show the derivative step does not always replace manual background removal. A complementary check is to compare DIVA's zero-crossing peak areas with those from a careful gold-standard baseline correction on the same plant data and look for systematic disagreement in conditions with high fluorescence.","tokens_in":24561,"feed_emoji":"🌱","tokens_out":10461,"duration_ms":97843,"temperature":0.7,"pith_summary":"This paper claims that a variational autoencoder fed only with the first derivative of raw Raman spectra can detect and quantify plant stress without any of the manual preprocessing that normally precedes Raman analysis: no fluorescence-background removal, no reference-intensity normalization, and no advance choice of which peaks matter. The claim matters because those manual steps vary from analyst to analyst, bias each experiment toward already known biomarkers, and force a separate pipeline for every stress, so an automated alternative would make plant health monitoring more reproducible and easier to deploy in the field. The authors apply the workflow, called DIVA, to Arabidopsis, Choy Sum, and Kai Lan plants under light stress, shade avoidance (including the photoreceptor mutants phyA-211 and phyB-9BC), high temperature, bacterial infection, and bacterial elicitors, and report that its latent-space clusters separate stress conditions and that its reconstructed spectra recover the expected stress markers of carotenoids, cellulose/lignin/proteins, and pectin, along with species-specific response trajectories. If the central claim is right, Raman-based stress detection in plants no longer requires a human expert to clean the spectra before a model can interpret them.","feed_headline":"Plant stress detected from raw Raman spectra, no cleanup needed","feed_subtitle":"Derivative spectra plus an autoencoder separate light, heat, and bacterial stress across crop species.","key_machinery":"The argument is carried by two objects working together. The first is the first-derivative transformation $D(\\tilde{\\nu}) = dI(\\tilde{\\nu})/d\\tilde{\\nu}$ of the raw Raman intensity, whose stated job is to suppress the slowly varying fluorescence background relative to the sharper Raman peaks so that baseline fitting, smoothing, and normalization can be skipped. The second is a variational autoencoder with a 64-neuron hidden layer and a 4-neuron latent layer, trained with an ELBO loss that combines mean-absolute-error reconstruction with KL divergence, producing a continuous and complete latent space in which spectra cluster by stress condition. The link between the two is the decoding step: the median of each cluster is decoded into a characteristic derivative spectrum, and a zero-crossing analysis of that reconstruction—specifically the crossings where the derivative changes from positive to negative—fixes the peak positions, while the area under the graph between neighboring crossings quantifies the biomolecular content at each wavenumber.","core_discovery":"The central claim, stated in the paper's own terms, is that native Raman spectra—fluorescence background included—can be analyzed end-to-end without manual preprocessing. The workflow first replaces each raw spectrum $I(\\tilde{\\nu})$ with its first derivative $D(\\tilde{\\nu}) = dI(\\tilde{\\nu})/d\\tilde{\\nu}$, on the argument that the slowly varying fluorescence background has a smaller derivative than the sharp Raman peaks, so baseline correction and normalization become unnecessary. The derivative spectra are then fed to a variational autoencoder whose two-dimensional latent space clusters spectra by plant condition, and the decoder reconstructs the characteristic derivative spectrum of each cluster from its median point. Raman peaks are located as the positive-to-negative zero-crossings of that reconstructed derivative, and the area under the graph at each peak provides a quantitative readout of the associated biomolecule without any prior hypothesis about which wavenumbers should respond. Across light, shade, heat, bacterial-infection, and elicitor experiments in three species and two Arabidopsis mutants, this unsupervised pipeline reports condition-specific latent-space separations and recovers the same core set of stress markers—carotenoids near 1521, 1151, and 1180 cm$^{-1}$; cellulose/lignin/protein near 1318 cm$^{-1}$; pectin near 742 cm$^{-1}$—while also flagging treatment-specific peaks such as 1001 versus 1178 cm$^{-1}$ in buffer- versus pathogen-infiltrated leaves and 843 cm$^{-1}$ under elf18 elicitation.","pith_inferences":["The derivative-plus-autoencoder trick may transfer to other spectroscopies with slowly varying backgrounds, such as near-infrared or fluorescence emission spectroscopy, where baseline removal is also a bottleneck; the paper develops it only for Raman.","Each case study trains its own model, so a single jointly trained DIVA that separates stress type, intensity, and time into distinct latent axes remains untested; if it worked, the latent space could become a general plant-health coordinate system rather than a per-experiment clustering.","The area-under-graph values are ranked rather than statistically tested; permuting or bootstrapping the per-group peak areas would turn 'significant peaks' into a formal significance claim, which the current workflow does not provide.","If latent-space geometry reflects physiology as faithfully as the phyB-9BC results suggest, DIVA could be used to phenotype unknown mutants by locating their spectra relative to known genotypes and predicting their stress trajectories before any phenotype measurement."],"forward_implications":["Plant stress detection no longer requires manual fluorescence-background removal, normalization, or pre-selected peak lists, so results do not depend on the analyst's choices.","One unsupervised pipeline spans abiotic and biotic stressors and multiple species and genotypes, replacing bespoke stress-by-stress workflows.","Molecular stress responses become visible before symptoms appear, as in the bacterial-infection study where latent separation and marker changes were detected at 24 hours with no visible signs of disease.","Latent-space trajectories expose species- and genotype-specific strategies: under heat, Arabidopsis declines, Choy Sum stabilizes after an early adjustment, and Kai Lan accumulates changes stepwise, while the phyB-9BC mutant clusters near its constitutively shade-avoiding phenotype across light conditions.","Peaks are found systematically from zero-crossings and quantified by area under the graph, so the same raw spectra yield the same analysis without investigator judgment."],"supporting_citations":[{"why":"Documents the manual preprocessing burden and user bias in Raman spectral analysis that DIVA claims to remove.","marker":"[19]"},{"why":"Represents the fluorescence-background-subtraction procedures that DIVA's derivative step is designed to bypass.","marker":"[21]"},{"why":"Supplies the shade-avoidance experimental system, the phyA-211 and phyB-9BC mutants, and the Raman band assignments for carotenoids, cell-wall polymers, and pectin used throughout.","marker":"[5]"},{"why":"Provides the elicitor (elf18/fig22) infiltration protocol for Arabidopsis and the reference Raman peaks used to label stress markers.","marker":"[15]"},{"why":"The source of the variational autoencoder formulation the workflow is built on.","marker":"[30]"},{"why":"The tutorial basis for the VAE architecture and ELBO training used in the encoder-decoder.","marker":"[31]"},{"why":"Standard reference for using Raman spectroscopy to characterize biological materials, grounding the spectral interpretation.","marker":"[14]"},{"why":"Source for the antagonistic roles of phytochrome A and B that motivate the mutant shade-avoidance expectations.","marker":"[27]"}],"fun_headline_variants":["AI reads raw Raman spectra to spot plant stress","No-prep Raman analysis IDs plant stress via deep learning","Deep learning detects plant stress from unprocessed Raman data","Autoencoder spots plant stress without spectral cleanup","Raman + AI: stress detection without preprocessing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the first derivative of the raw Raman spectrum suppresses the slowly varying fluorescence background strongly enough that baseline correction and normalization can be dropped, and the paper does not test this on spectra with steep or strongly structured fluorescence.","fun_headline_variants_meta":{"raw":{"variants":["AI reads raw Raman spectra to spot plant stress","No-prep Raman analysis IDs plant stress via deep learning","Deep learning detects plant stress from unprocessed Raman data","Autoencoder spots plant stress without spectral cleanup","Raman + AI: stress detection without preprocessing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1457,"prompt_tokens":1049,"completion_tokens":408,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":665,"completion_tokens_details":{"reasoning_tokens":336}},"tokens_in":665,"tokens_out":408,"duration_ms":4563,"temperature":1.0,"reasoning_tokens":336,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:23:48.265754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate or collect Raman spectra with known peaks embedded in a steep, structured fluorescence background, such as an exponential or sharply sloped baseline, run them through the DIVA workflow, and check whether the detected peak positions and relative areas match the known inputs; spurious zero-crossing peaks, shifted positions, or noise-dominated reconstructions would show the derivative step does not always replace manual background removal. A complementary check is to compare DIVA's zero-crossing peak areas with those from a careful gold-standard baseline correction on the same plant data and look for systematic disagreement in conditions with high fluorescence.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the manual preprocessing burden and user bias in Raman spectral analysis that DIVA claims to remove."},{"cited_title":"& Stoddart, P","cited_arxiv_id":null,"evidence_quote":"Represents the fluorescence-background-subtraction procedures that DIVA's derivative step is designed to bypass."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the shade-avoidance experimental system, the phyA-211 and phyB-9BC mutants, and the Raman band assignments for carotenoids, cell-wall polymers, and pectin used throughout."},{"cited_title":"J., Seo, J","cited_arxiv_id":null,"evidence_quote":"Provides the elicitor (elf18/fig22) infiltration protocol for Arabidopsis and the reference Raman peaks used to label stress markers."},{"cited_title":"P ., Welling, M","cited_arxiv_id":null,"evidence_quote":"The source of the variational autoencoder formulation the workflow is built on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Standard reference for using Raman spectroscopy to characterize biological materials, grounding the spectral interpretation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source for the antagonistic roles of phytochrome A and B that motivate the mutant shade-avoidance expectations."}],"review_version":1}