Pith. sign in

REVIEW 4 major objections 7 minor 49 references

Brightfield images predict Raman spectra at 98% similarity

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-09 03:36 UTC pith:ZCKCZRRW

load-bearing objection Novel brightfield-to-Raman translation with suggestive but insufficiently validated results; missing mean-spectrum baseline is the key gap. the 4 major comments →

arxiv 2607.07651 v1 pith:ZCKCZRRW submitted 2026-07-08 physics.optics physics.app-phphysics.comp-phstat.CO

Pic2Spec: Generative Modeling Reconstructs Single Cell Raman Fingerprints from Brightfield Images

classification physics.optics physics.app-phphysics.comp-phstat.CO
keywords Raman spectroscopybrightfield microscopyvariational autoencodercross-modal translationsingle-cell analysisgenerative modellabel-free phenotypingspectral reconstruction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Pic2Spec claims that the morphological information in a standard brightfield microscopy image is sufficient to computationally reconstruct a single cell's Raman spectrum — its label-free biochemical fingerprint — without any spectroscopic hardware. The authors train a dual-decoder variational autoencoder on paired brightfield-Raman measurements, where a shared latent representation simultaneously reconstructs the input image and decodes a 573-point Raman spectrum. Across Jurkat T cells, primary B cells, and E. coli, generated spectra match measured spectra at ~98% cosine similarity and ~95% Pearson correlation. In bacterial systems, the generated spectra discriminate GFP-expressing from non-expressing E. coli at 88% accuracy, compared to 68% for image-only classification, suggesting the model recovers molecular information not directly visible in morphology. The authors argue this constitutes the first demonstration of chemically informative virtual molecular fingerprints inferred purely from brightfield contrast.

Core claim

The central object is a dual-decoder variational autoencoder architecture in which a single encoder maps a 96×96 brightfield image into a 100-dimensional Gaussian latent space, from which two decoders jointly reconstruct the image and predict the Raman spectrum. The image-reconstruction branch acts as a structural anchor, preventing the latent space from drifting toward spectral-only features untethered from the observed cell morphology. The spectral decoder uses 1D convolutions with a composite loss combining mean squared error, cosine distance, derivative alignment, and peak-ratio preservation. The authors show this joint architecture preserves cell-to-cell spectral variability better than

What carries the argument

A dual-decoder VAE with a shared latent space linking brightfield morphology to Raman spectral output, trained with a four-component spectral loss (MSE, cosine, derivative, peak-ratio) and KL annealing.

Load-bearing premise

The evaluation rests on very small numbers of unique test cells (roughly 20–45 original cells per system after an 80:10:10 split, expanded 6-fold by augmentation), and the test spectra themselves are augmented copies of measurements from the same experimental session rather than independently acquired spectra, making it difficult to distinguish genuine cross-modal learning from memorization of session-specific correlations.

What would settle it

Train Pic2Spec on one experimental session's paired data, then test on paired brightfield-Raman data acquired in a separate session with different cell preparations; if spectral similarity drops sharply, the model has learned session-specific correlations rather than a generalizable morphology-to-biochemistry mapping.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the image-to-spectrum mapping generalizes across microscope platforms and cell states, any laboratory with a brightfield microscope could perform label-free biochemical phenotyping without purchasing Raman hardware.
  • The framework could extend to longitudinal live-cell monitoring, where repeated Raman measurements are impractical due to phototoxicity or throughput, by inferring molecular profiles from time-lapse brightfield images alone.
  • The latent-space disentanglement of intensity from composition suggests that controlled latent perturbation could be used to explore how morphological changes map to specific biochemical shifts, enabling hypothesis generation about morphology-composition relationships.
  • If the approach extends to clinical samples (blood, tissue aspirates), it could enable rapid pathogen identification or cell-state screening at point-of-care settings where Raman instrumentation is unavailable.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The high spectral similarity metrics may be partly explained by the relative homogeneity of Raman spectra within a given cell type; if the model learns to output a cell-type-conditioned mean spectrum with modest image-driven variation, cosine similarity and Pearson correlation would remain high even without a deep morphology-to-biochemistry mapping. The 20-point improvement in GFP classification o
  • The saliency analysis showing spatially structured attribution within cell footprints is suggestive but not conclusive: a model that has learned session-specific correlations between image artifacts and spectral features could also produce cell-localized saliency maps.
  • A decisive test would be to acquire paired brightfield-Raman data on one instrument, train the model, then evaluate on cells imaged on a different microscope or from a different culture batch, with independently acquired Raman spectra as ground truth rather than augmented copies of training-session measurements.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This manuscript introduces Pic2Spec, a dual-decoder variational autoencoder that predicts single-cell Raman spectra from brightfield microscopy images. The model is trained and evaluated on paired image-spectrum datasets from Jurkat T cells, primary B cells, and E. coli (including GFP+ and GFP- strains). The authors report 98% cosine similarity and ~95% Pearson correlation between generated and measured spectra, and show that generated spectra discriminate GFP+ from GFP- E. coli at ~88% accuracy, outperforming an image-only baseline by ~20 percentage points. Architectural comparisons among three model variants, saliency mapping, and latent-space perturbation analyses are provided to argue that the learned mapping is structured and biochemically grounded. The central claim is that this constitutes the first demonstration of chemically informative virtual molecular fingerprints inferred from brightfield contrast alone.

Significance. If the central claim holds, this work would represent a meaningful advance in computational spectroscopy, potentially enabling molecular profiling from ubiquitous microscopy platforms. The dual-decoder VAE architecture, the multi-component spectral loss, and the latent-space interpretability analyses represent a thoughtful engineering effort. The GFP classification result, where generated spectra outperform image-only analysis, is the most compelling piece of evidence for image-conditioned biochemical inference. The saliency and latent perturbation analyses are commendable attempts at mechanistic interpretation.

major comments (4)
  1. No mean-spectrum baseline is reported anywhere in the manuscript. This is load-bearing for the central claim of 'cell-dependent spectra prediction' (Results, paragraph 4: 'Pic2Spec captures cell-resolved spectral structure... rather than merely reproducing population-level averages'). Raman spectra within a cell type share highly conserved major vibrational bands; a model that simply outputs the cell-type mean training spectrum would likely achieve very high cosine similarity and Pearson correlation. The paper's own architectural comparison provides indirect evidence of this risk: the Enc-Dec model is described as showing 'regression-to-the-mean behavior' with 'markedly compressed' variability (Fig. 4D-E), yet it still achieves cosine similarity and Pearson correlation values comparable to Pic2Spec (Fig. 4A-C(iii-iv)). This suggests that global similarity metrics are insensitive to cell-
  2. The effective test set sizes are very small after de-augmentation. Methods §5-6 state that each cell is augmented 6-fold (5 geometric transforms + original) and that augmented variants of the same cell are confined to a single split. With 206 T-cell and 209 B-cell pairs at an 80:10:10 split, the test sets of ~123 T cells and ~125 B cells derive from approximately 20-21 unique cells each. The bacterial test set of 270 derives from ~45 unique cells. These sample sizes are small for deep learning evaluation and limit the statistical reliability of the reported metrics and their confidence intervals. The manuscript should report the number of unique (de-augmented) test cells alongside the augmented counts and discuss the implications for generalization.
  3. Test spectra are augmented versions of original measurements (intensity scaling ±10%, Gaussian noise 0.5-2%, peak broadening 0-3 cm^-1; Methods §5), not independently acquired data. While the augmentation is designed to simulate experimental variability, the model is evaluated against perturbed copies of spectra from the same experimental session on the same instrument. This makes it difficult to distinguish genuine cross-modal learning from memorization of session-specific correlations. The manuscript should acknowledge this limitation explicitly and discuss how performance might change on independently acquired spectra from different sessions or instruments.
  4. The spectral loss function includes a peak-ratio term (L_ratio, Methods §7, Eq. for L_ratio) that constrains selected peak-intensity ratios (I_1008/I_1120, I_720/I_740, I_995/I_1010). The same ratios are then evaluated in Fig. 5B to demonstrate that 'Pic2Spec retained relative intensity structure between functionally related peaks rather than reproducing the average spectral shape.' Since these ratios are directly optimized during training, their preservation in generated spectra is partly expected by construction and does not independently demonstrate that the model learns a generalizable morphological-to-biochemical mapping. The authors should clarify this circularity and evaluate peak ratios that were not included in the loss function.
minor comments (7)
  1. The abstract states '98% cosine similarity and Pearson correlations of ~95%' without specifying that these are median values; the main text (Fig. 2C) reports median Pearson r = 0.94 [0.90-0.95] for T cells, which is slightly below ~95%. Clarify.
  2. Fig. 2C caption states 'n=123 T cells, n=125 B cells' but does not clarify that these are augmented counts. Adding the unique cell count would improve transparency.
  3. The notation in Methods §7 for the spectral loss uses both 'K' and 'L' for spectral length in different equations (L_MSE uses K, while the text later refers to L=573). Standardize.
  4. The SAM metric is defined as 'R_S(y, ŷ)' in Methods §8 but the symbol is unconventional; consider using θ or SAM to avoid confusion with Pearson's r.
  5. Fig. 5E caption refers to 'global intensity effect' and 'composition-sensitive (peak-ratio) effect' but the axes labels in the figure should match the definitions in Methods §12 (I_k and C_k) for clarity.
  6. The claim 'first demonstration of chemically informative virtual molecular fingerprints inferred purely from brightfield contrast' appears in both the abstract and significance statement. Given that the evaluation is limited to same-session data with augmented test spectra, this claim should be tempered or qualified.
  7. References 4 and 19 in the Supporting Information cite future dates (2025, 2026); verify these are correctly published or mark as 'in press.'

Simulated Author's Rebuttal

4 responses · 2 unresolved

We thank the referee for a careful and substantive review. The comments identify legitimate concerns regarding baseline comparisons, effective sample sizes, spectral augmentation in the test set, and potential circularity in peak-ratio evaluation. We address each point below and describe concrete revisions we will make.

read point-by-point responses
  1. Referee: No mean-spectrum baseline is reported. A model outputting the cell-type mean training spectrum would likely achieve high cosine similarity and Pearson correlation, and the Enc-Dec model shows regression-to-the-mean yet achieves comparable global metrics. This is load-bearing for the claim of cell-resolved spectral structure.

    Authors: The referee is correct that a mean-spectrum baseline is essential and its absence is a significant gap. We will add this baseline in revision. We will compute the cell-type mean training spectrum and evaluate it against the held-out test set using all four metrics (RMSE, cosine similarity, Pearson correlation, SAM) for each cell system (T cells, B cells, bacteria, GFP+, GFP-). We expect the mean baseline to achieve high cosine similarity and Pearson correlation, which would confirm the referee's concern that global metrics alone are insufficient to demonstrate cell-resolved prediction. This is precisely why we included the band-area and peak-height variability analyses (Fig. 4D-E) and the single-cell spectral heterogeneity figures (Figs. S5, S7, S9): these show that Pic2Spec preserves cell-to-cell variability that the Enc-Dec model compresses. However, we agree that a direct quantitative comparison against the mean baseline strengthens this argument and should have been included. We will add it and will also report per-cell residual variance metrics (e.g., the standard deviation of residuals across cells) that are more sensitive to cell-resolved structure than global similarity metrics. We will temper the claim about 'cell-resolved spectral structure' to explicitly acknowledge that global similarity metrics cannot distinguish cell-resolved prediction from mean reproduction, and that our evidence for cell-resolved structure rests on the variability preservation analyses rather than the aggregate metrics. revision: yes

  2. Referee: Effective test set sizes are very small after de-augmentation (~20-21 unique cells for T and B cells, ~45 for bacteria). These sample sizes limit statistical reliability of reported metrics and confidence intervals.

    Authors: The referee's arithmetic is correct. With 6-fold augmentation and 80:10:10 splits, the test sets of 123 T-cell, 125 B-cell, and 270 bacterial spectra derive from approximately 20-21, 20-21, and 45 unique cells, respectively. We agree these are small for deep learning evaluation and that this limitation should be transparently reported. In revision, we will (1) report the number of unique de-augmented test cells alongside augmented counts in all relevant figure captions and the Methods, (2) add a discussion of the implications for generalization, and (3) note that confidence intervals reported on augmented test samples do not reflect independent biological replicates. We acknowledge that this is a genuine limitation of the current study that cannot be fully resolved without additional data collection. We will state this explicitly as a scope limitation and identify larger-scale data acquisition as a priority for future work. We respectfully note that the augmented test samples do represent non-overlapping cells (augmented variants of the same cell are confined to one split), so there is no train-test leakage; the issue is statistical power, not data contamination. revision: yes

  3. Referee: Test spectra are augmented versions of original measurements (intensity scaling, noise, peak broadening), not independently acquired data. This makes it difficult to distinguish genuine cross-modal learning from memorization of session-specific correlations.

    Authors: This is a fair and important concern. The spectral augmentation (intensity scaling ±10%, Gaussian noise 0.5-2%, peak broadening 0-3 cm^-1) is applied to spectra from the same experimental session and instrument, so the test set does not represent fully independent spectral measurements. We agree that this should be acknowledged as a limitation. In revision, we will add an explicit discussion of this point in the manuscript, noting that (1) the test spectra are perturbed copies of same-session measurements rather than independently acquired data, (2) performance on truly independent spectra from different sessions or instruments may differ due to session-specific correlations, calibration drift, or instrument variability, and (3) cross-session and cross-instrument validation is needed to establish robustness to domain shift. We note that the brightfield images in the test set are geometrically augmented (rotations, flips) rather than intensity-perturbed, and the image-to-spectrum mapping must still generalize across orientations. However, the referee's core point stands: the spectral targets share session-specific characteristics with training data, and we cannot rule out that the model exploits some session-specific correlations. We will state this limitation clearly and identify cross-session validation as a critical next step. We cannot fully resolve this concern with the current dataset. revision: yes

  4. Referee: The spectral loss includes L_ratio constraining peak ratios (I_1008/I_1120, I_720/I_740, I_995/I_1010), and the same ratios are evaluated in Fig. 5B. This is circular: their preservation is partly expected by construction.

    Authors: The referee is correct that evaluating peak ratios that are directly optimized in the loss function is circular and does not independently demonstrate generalization. We will address this in two ways. First, we will explicitly acknowledge in the manuscript that the three ratios shown in Fig. 5B (I_1008/I_1120, I_720/I_740, I_995/I_1010) are the same ratios included in L_ratio, and that their preservation is therefore partly expected by construction. Second, we will evaluate additional peak ratios that were NOT included in the loss function. Specifically, we will compute ratios such as I_1450/I_1660 (CH2/Amide I), I_1003/I_1200 (phenylalanine/amide III), and I_785/I_1095 (nucleic acid/phosphate) between true and generated spectra, and report JS divergence for these non-optimized ratios. If these ratios are also preserved, it would provide independent evidence that the model learns a generalizable morphological-to-biochemical mapping rather than merely satisfying the loss constraints. If they are not preserved, we will report that honestly and adjust our claims accordingly. We will revise the Fig. 5B analysis and accompanying text to separate optimized from non-optimized ratios and to clarify which conclusions rest on independent evidence. revision: yes

standing simulated objections not resolved
  • The small number of unique cells (20-45 per test set) is a fundamental limitation of the current dataset that cannot be resolved without new data collection. We can acknowledge it transparently but cannot increase the sample sizes in revision.
  • Cross-session and cross-instrument validation cannot be performed with existing data, as all measurements were acquired on a single instrument in single sessions per cell type. We can acknowledge this as a limitation but cannot address it within the scope of this revision.

Circularity Check

0 steps flagged

No significant circularity; minor overlap between training loss components and evaluation metrics, but central claims supported by independent evaluation on held-out data.

full rationale

Pic2Spec is a supervised generative model trained on paired brightfield-Raman data and evaluated on held-out test cells. The spectral reconstruction loss includes cosine similarity (λ=0.2) and peak-ratio (λ=0.005) terms, which overlap with some reported evaluation metrics (cosine similarity in Fig. 2C/G, peak ratios in Fig. 5B). However, this is standard supervised ML practice: the model is evaluated on non-overlapping held-out cells, and high metrics on test data are not guaranteed by the training objective. The peak-ratio loss uses an unspecified set P of peak pairs with a very small weight (0.005), and the peak-ratio evaluation uses JS divergence at the distribution level rather than per-sample loss. The paper's central claim — that generated spectra discriminate GFP+ from GFP- E. coli at 88% accuracy — is supported by an independent downstream classification task not tied to the spectral reconstruction loss. The sole self-citation (Ref. 19, SpectroGen by co-author Tadesse) is contextual and not load-bearing. No derivation step reduces to its inputs by construction.

Axiom & Free-Parameter Ledger

6 free parameters · 3 axioms · 0 invented entities

No new physical entities, particles, forces, or dimensions are introduced. The dual-decoder VAE architecture is a composition of existing neural network components.

free parameters (6)
  • λ_img (image reconstruction loss weight) = 0.5
    Chosen by the authors; controls the trade-off between image reconstruction and spectral generation in the dual-decoder VAE.
  • λ_MSE, λ_cos, λ_deriv, λ_ratio (spectral loss weights) = 1.0, 0.2, 0.05, 0.005
    Hand-set weights for the composite spectral loss; not derived from first principles or cross-validated over a range.
  • Latent dimensionality = 100
    Chosen for the latent space dimension; no systematic search reported.
  • Convolutional filter sizes (8, 16, 4, 4) = 8, 16, 4, 4
    Stated as chosen based on literature; no ablation over architecture depth or width reported.
  • β-annealing schedule = Not specified numerically
    Time-dependent KL weight schedule is mentioned but the specific annealing curve is not provided.
  • Spectral augmentation parameters = shift ±1.5 cm⁻¹, intensity ±10%, noise 0.5-2%, broadening 0-3 cm⁻¹
    Hand-chosen ranges for spectral augmentation applied to training and test data.
axioms (3)
  • domain assumption There exists a learnable mapping from brightfield image morphology to Raman spectral structure at the single-cell level.
    This is the foundational premise of the entire paper. It is plausible but unproven; the strength of the morphological-biochemical correlation determines the ceiling of achievable performance.
  • domain assumption Brightfield images do not contain fluorescence or other non-morphological chemical information.
    The paper frames brightfield as encoding only refractive index and thickness. If the white-light LED excites GFP or other autofluorescence, this assumption is violated and the model may be exploiting fluorescence rather than morphology.
  • ad hoc to paper Augmented spectra are valid proxies for independently acquired measurements.
    The evaluation uses augmented (perturbed) versions of original spectra as test targets. This assumes that performance on perturbed copies generalizes to genuinely new measurements, which is not tested.

pith-pipeline@v1.1.0-glm · 26192 in / 5488 out tokens · 576963 ms · 2026-07-09T03:36:41.384756+00:00 · methodology

0 comments
read the original abstract

Single-cell molecular characterization remains a bottleneck in scalable biological analysis because of labeling requirements, limited multiplexing, and reagents that perturb physiology. Raman spectroscopy addresses these limits by providing chemically specific, label-free vibrational fingerprints, but long acquisition times and specialized instruments restrict high-throughput use. Here, we overcome this barrier by showing that spectral fingerprints can be reconstructed from brightfield microscopy using generative modeling. We introduce Pic2Spec, a framework that learns a shared latent biochemical representation linking image morphology to vibrational spectral structure, enabling virtual Raman spectroscopy without hardware. We validate Pic2Spec across mammalian and bacterial cells, generating high-fidelity spectra that reproduce measured Raman fingerprints with 98% cosine similarity and Pearson correlations of ~95%, while preserving biochemical peaks and population distributions. Beyond spectral similarity, Pic2Spec provides molecular-level resolution in bacterial systems: generated spectra discriminate mutation-driven transgenic states and predict GFP expression with accuracy approaching true Raman measurements, outperforming conventional image analysis by 20%. These findings establish Pic2Spec as a first demonstration of chemically informative virtual molecular fingerprinting from brightfield images, complementing slow, hardware-intensive spectroscopy with computational inference. By redefining microscopy as an inference-enabled molecular profiling platform, Pic2Spec democratizes label-free biochemical phenotyping and overcomes the hardware and time constraints that have confined spectroscopy to specialized laboratories. This enables high-throughput molecular analysis for clinical diagnostics, screening, and monitoring at the scale and accessibility of standard microscopy.

Figures

Figures reproduced from arXiv: 2607.07651 by Amit Kumar Bhuyan, Loza F. Tadesse, Srilakshmi Premachandran.

Figure 1
Figure 1. Figure 1: Schematic representation of Pic2Spec framework: Single [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Pic2Spec recovers single-cell Raman spectra from brightfield images in immune cells. Left panel: Jurkat T cells. Right panel: Primary B cells. A,E. Representative single-cell BF images used as model input. B,F. Generated Raman spectra with the corresponding measured spectra across the fingerprint region (600-1800 cm−1 ). Shaded regions indicate ± 1 standard deviation (s.d.) C,G. Distributions of spectral s… view at source ↗
Figure 3
Figure 3. Figure 3: Pic2Spec generalizes to bacterial cells and preserves phenotype-discriminative spectral information. Spectral reconstruction performance for GFP− and GFP+ E. coli cells. (A), (B) Generated Raman spectra with the corresponding measured spectra for GFP− and GFP+ cells, respectively. Insets show representative BF and fluorescence images used to assign GFP class labels. Spectra show mean ± 1 s.d. Across n= 132… view at source ↗
Figure 4
Figure 4. Figure 4: Joint image-spectrum latent learning improves spectral realism across distinct cell systems: Comparison of three image-to-spectrum translation architectures: a direct cross-modal encoder-decoder (Enc-Dec), the proposed dual-decoder variational model (Dualdec VAE / Pic2Spec), and a two-step latent-translation variational model (Lat-Trans VAE). A-C, Performance across three cellular systems: A) pooled bacter… view at source ↗
Figure 5
Figure 5. Figure 5: Fig.5. Band- [PITH_FULL_IMAGE:figures/full_fig_p018_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

49 extracted references · 49 canonical work pages · 1 internal anchor

  1. [1]

    Chen, B. et al. Label-free live cell recognition and tracking for biological discoveries and translational applications. npj Imaging 2, 41 (2024)

  2. [2]

    McKinnon, K. M. Flow Cytometry: An Overview. Current Protocols in Immunology 120, 5.1.1-5.1.11 (2018)

  3. [3]

    P., Ostafe, R., Iyengar, S

    Robinson, J. P., Ostafe, R., Iyengar, S. N., Rajwa, B. & Fischer, R. Flow Cytometry: The Next Revolution. Cells 12, 1875 (2023)

  4. [4]

    L., Atoyan, J., Cai, A

    Hemme, C. L., Atoyan, J., Cai, A. & Liu, C. Challenges and Opportunities in Multi -Omics Data Acquisition and Analysis: Toward Integrative Solutions. Biomolecules 16, 271 (2026)

  5. [5]

    Record, C. J. & Reilly, M. M. Lessons and pitfalls of whole genome sequencing. Practical Neurology 24, 263–274 (2024)

  6. [6]

    Kobayashi-Kirschvink, K. J. et al. Prediction of single-cell RNA expression profiles in live cells by Raman microscopy with Raman2RNA. Nat Biotechnol 42, 1726–1734 (2024)

  7. [7]

    Kamei, K. F. & Wakamoto, Y. Live-cell omics with Raman spectroscopy. Microscopy (Oxf) 74, 189–200 (2025)

  8. [8]

    & Smith, N

    Pavillon, N. & Smith, N. I. Non-invasive monitoring of T cell differentiation through Raman spectroscopy. Sci Rep 13, 3129 (2023)

  9. [9]

    Pavillon, N. et al. Non-invasive detection of regulatory T cells with Raman spectroscopy. Sci Rep 14, 14025 (2024)

  10. [10]

    Zhang, Y. et al. From Genotype to Phenotype: Raman Spectroscopy and Machine Learning for Label-Free Single-Cell Analysis. ACS Nano 18, 18101–18117 (2024)

  11. [11]

    Li, Y. et al. Rapid culture-free diagnosis of clinical pathogens via integrated microfluidic - Raman micro-spectroscopy. Nat Commun 17, 283 (2025)

  12. [12]

    & Stevens, M

    Fernández-Galiana, Á., Bibikova, O., Vilms Pedersen, S. & Stevens, M. M. Fundamentals and Applications of Raman- Based Techniques for the Design and Development of Active Biomedical Materials. Advanced Materials 36, 2210807 (2024)

  13. [13]

    & Mahadevan -Jansen, A

    Pence, I. & Mahadevan -Jansen, A. Clinical instrumentation and applications of Raman spectroscopy. Chem Soc Rev 45, 1958–1979 (2016)

  14. [14]

    Harrison, P. J. et al. Evaluating the utility of brightfield image data for mechanism of action prediction. PLoS Comput Biol 19, e1011323 (2023)

  15. [15]

    K., Therkildsen, M

    Skrivergaard, S., Rasmussen, M. K., Therkildsen, M. & Young, J. F. High- Throughput Label-Free Continuous Quantification of Muscle Stem Cell Proliferation and Myogenic Differentiation. Stem Cell Rev Rep 21, 2103–2120 (2025)

  16. [16]

    & Teichmann, S

    Schuster, V., Dann, E., Krogh, A. & Teichmann, S. A. multiDGD: A versatile deep generative model for multi-omics data. Nat Commun 15, 10031 (2024)

  17. [17]

    Radhakrishnan, A. et al. Cross-modal autoencoder framework learns holistic representations of cardiovascular state. Nat Commun 14, 2436 (2023)

  18. [18]

    & Torr, P

    Shi, Y., N, S., Paige, B. & Torr, P. Variational Mixture -of-Experts Autoencoders for Multi- Modal Deep Generative Models. in Advances in Neural Information Processing Systems vol. 32 (Curran Associates, Inc., 2019)

  19. [19]

    & Tadesse, L

    Zhu, Y. & Tadesse, L. F. SpectroGen: A physically informed generative artificial intelligence for accelerated cross-modality spectroscopic materials characterization. Matter 9, (2026)

  20. [21]

    Y., Kalinin, S

    Yaman, M. Y., Kalinin, S. V., Guye, K. N., Ginger, D. S. & Ziatdinov, M. Learning and Predicting Photonic Responses of Plasmonic Nanoparticle Assemblies via Dual Variational Autoencoders. Small 19, (2023)

  21. [22]

    Haque, M. I. U. et al. Deep learning- driven super -resolution in Raman hyperspectral imaging: Efficient high- resolution reconstruction from low -resolution data. Appl. Phys. Lett. 125, (2024)

  22. [23]

    Georgiev, D. et al. Hyperspectral unmixing for Raman spectroscopy via physics - constrained autoencoders. Proceedings of the National Academy of Sciences 121 , e2407439121 (2024)

  23. [24]

    He, H. et al. Noise learning of instruments for high- contrast, high -resolution and fast hyperspectral microscopy and nanoscopy. Nat Commun 15, 754 (2024)

  24. [25]

    Pathak, P. et al. Spectral Similarity Measures for In Vivo Human Tissue Discrimination Based on Hyperspectral Imaging. Diagnostics (Basel) 13, 195 (2023)

  25. [26]

    Ho, C.-S. et al. Rapid identification of pathogenic bacteria using Raman spectroscopy and deep learning. Nat Commun 10, 4927 (2019)

  26. [27]

    Evans, T. D. & Zhang, F. Bacterial Metabolic Heterogeneity: Origins and Applications in Engineering and Infectious Disease. Curr Opin Biotechnol 64, 183–189 (2020)

  27. [28]

    & Skandamis, P

    Arvaniti, M. & Skandamis, P. N. Defining bacterial heterogeneity and dormancy with the parallel use of single-cell and population level approaches. Current Opinion in Food Science 44, 100808 (2022)

  28. [29]

    Uemura, K. et al. Rapid and Integrated Bacterial Evolution Analysis unveils gene mutations and clinical risk of Klebsiella pneumoniae. Nat Commun 16, 2917 (2025). Supporting Information Pic2Spec: Generative Modeling Reconstructs Single Cell Raman Fingerprints from Brightfield Images Srilakshmi Premachandran1, Amit Kumar Bhuyan1, Loza F. Tadesse1,2,3* 1Dep...

  29. [30]

    coli strain Seattle 1946 cells (#25922 and #25922GFP for GFP− and GFP+ cells, respectively) were purchased from American Type Culture Collection (ATCC)

    Cell culture E. coli strain Seattle 1946 cells (#25922 and #25922GFP for GFP− and GFP+ cells, respectively) were purchased from American Type Culture Collection (ATCC). Cells stored as frozen glycerol stock were streaked onto polystyrene plates with solidified LB agar media (Teknova, Cat. #L1100). Cells were incubated overnight at 37 ℃. A single bacterial...

  30. [31]

    Images of single cells were collected using a 100x objective (Zeiss EC Epiplan- Neofluar, NA = 0.9) at 50% illumination intensity

    Acquisition of BF Images and Raman spectra of single bacterial cells BF images were acquired using a Witec Alpha 300 high-intensity white-light LED lamp for Köhler illumination, providing uniform illumination of the sample. Images of single cells were collected using a 100x objective (Zeiss EC Epiplan- Neofluar, NA = 0.9) at 50% illumination intensity. Th...

  31. [32]

    Each raw image was loaded and smoothed using a Gaussian filter to suppress high-frequency noise while preserving cell boundaries

    BF image preprocessing The collected BF images were processed using the Cellpose deep-learning cellular segmentation model to filter single cells morphologically1. Each raw image was loaded and smoothed using a Gaussian filter to suppress high-frequency noise while preserving cell boundaries. Segmentation of single cells was performed using the pretrained...

  32. [33]

    To correct for cosmic ray spikes, each raw spectrum was passed through a 1D median filter, and an intensity vector was computed using a kernel size of 5 points

    Raman spectral preprocessing Single-cell Raman spectra were preprocessed using a custom Python pipeline for cosmic ray removal, truncation, baseline correction, smoothing, and normalization. To correct for cosmic ray spikes, each raw spectrum was passed through a 1D median filter, and an intensity vector was computed using a kernel size of 5 points. The d...

  33. [34]

    Each input image was augmented using five geometric transformations, providing invar iance to cell orientation

    Data Augmentation To increase the effective size of training data, the preprocessed single cell images of bacteria were subjected to geometric augmentation using the Pillow imaging library. Each input image was augmented using five geometric transformations, providing invar iance to cell orientation. This included horizontal and vertical flips, 90, 180, a...

  34. [35]

    coli and E

    Data stratification The image and spectral dataset consists of data acquired from 2 bacterial classes (E. coli and E. coli GFP+). Data stratification and splitting were performed to ensure class balance and prevent any augmentation leakage between the training, validation, and test data. An 80-10-10 split was performed for training, validation, and test d...

  35. [36]

    Schematic comparison of generative architectures for image-to-Raman translation

    Model descriptions Scheme S1. Schematic comparison of generative architectures for image-to-Raman translation. The three model designs evaluated in this study are illustrated. Top, Dual-decoder VAE: a shared image encoder maps the input brightfield image into a common latent space that is decoded through two branches to jointly reconstruct the input image...

  36. [37]

    This representation was flattened to 144 units and projected to the latent mean and log- variance vectors

    Spatial downsampling was performed using 2×2 max pooling between convolutional blocks, yielding a final encoder feature map of size 6×6×4. This representation was flattened to 144 units and projected to the latent mean and log- variance vectors. Latent sampling was performed using the reparameterization trick, 𝑧𝑧= 𝝁𝝁 + 𝑒𝑒𝑥𝑥𝑒𝑒( 1 2 𝑙𝑙𝑙𝑙𝑙𝑙𝞼𝞼2) ⊙ 𝛆𝛆 , with 𝛆...

  37. [38]

    The resulting feature map was flattened and projected to a 100- dimensional latent representation through a fully connected layer

    Each convolution was followed by 2 × 2 max pooling, progressively reducing the spatial resolution from 96 × 96 to 6 × 6. The resulting feature map was flattened and projected to a 100- dimensional latent representation through a fully connected layer. The spectral decoder transformed the latent representation into a Raman spectrum of length 573. The laten...

  38. [39]

    All metrics were computed independently for each spectrum over the full spectral length

    Evaluation metrics Model performance was evaluated by comparing each predicted Raman spectrum variable( 𝑦𝑦�) with its corresponding ground-truth spectrum variable(y) using four complementary metrics: root mean squared error (RMSE), cosine similarity, Pearson correlation coefficient, and spectral angle mapper (SAM). All metrics were computed independently ...

  39. [40]

    For each spectrum, peak detection was performed within the Raman window using a prominence- based approach implemented in SciPy

    Peak characteristics analysis Peak-level spectral characteristics were quantified from each Raman spectrum to assess whether generated spectra preserved key local features of the true measurements. For each spectrum, peak detection was performed within the Raman window using a prominence- based approach implemented in SciPy. Peaks were required to exceed ...

  40. [41]

    Saliency Mapping Gradient-based saliency maps were generated to identify image regions contributing to spectral prediction. For a given input BF image 𝑥𝑥, the trained model produced a predicted spectrum 𝑦𝑦�, and a scalar target s( 𝑦𝑦�) was defined as the summed predicted intensity within a selected spectral band. Band limits specified in wavenumber space ...

  41. [42]

    coli was performed to determine whether generated Raman spectra preserved class-relevant molecular information

    Classification models Binary classification of GFP+ and GFP− E. coli was performed to determine whether generated Raman spectra preserved class-relevant molecular information. A common held-out test set was maintained across all evaluation settings to ensure direct comparability. Three matched testing conditions were used: Test 1, classification from expe...

  42. [43]

    This approach examines how controlled perturbations of latent coordinates propagate through the spectral decoder and alter the generated Raman spectra

    Interpretability Analyses To interrogate how the learned representation organizes spectral variability and to determine whether individual latent variables encode physically meaningful transformations, systematic perturbation-based analysis of the latent space was performed. This approach examines how controlled perturbations of latent coordinates propaga...

  43. [44]

    & Pachitariu, M

    Stringer, C., Wang, T., Michaelos, M. & Pachitariu, M. Cellpose: a generalist algorithm for cellular segmentation. Nat Methods 18, 100–106 (2021)

  44. [45]

    Probabilistic Latent Variable Models: Principles and Foundations for Modern Generative AI

    Chen, T. Probabilistic Latent Variable Models: Principles and Foundations for Modern Generative AI. SSRN Scholarly Paper at https://doi.org/10.2139/ssrn.5244929 (2025)

  45. [46]

    Variational Encoder-Decoders for Learning Latent Representations of Physical Systems

    Venkatasubramanian, S. & Barajas-Solano, D. A. Variational Encoder-Decoders for Learning Latent Representations of Physical Systems. arXiv.org https://arxiv.org/abs/2412.05175v1 (2024)

  46. [47]

    S., Guevarra, D., Newhouse, P

    Stein, H. S., Guevarra, D., Newhouse, P. F., Soedarmadji, E. & Gregoire, J. M. Machine learning of optical properties of materials – predicting spectra from images and images from spectra. Chem. Sci. 10, 47–55 (2018)

  47. [48]

    & Mondal, J

    Adhikari, S. & Mondal, J. Elucidating Protein Dynamics through the Optimal Annealing of Variational Autoencoders. J. Chem. Theory Comput. 21, 6367–6379 (2025)

  48. [49]

    & Mahmood, A

    Wei, R. & Mahmood, A. Recent Advances in Variational Autoencoders With Representation Learning for Biomedical Informatics: A Survey. IEEE Access 9, 4939–4956 (2021)

  49. [50]

    Li, P., Pei, Y. & Li, J. A comprehensive survey on design and application of autoencoder in deep learning. Applied Soft Computing 138, 110176 (2023)