{"id":"b28ddd62-5a8d-475d-b829-517eaafa8e01","arxiv_id":"2505.08123","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"JSover jointly reconstructs material volume fractions and the X-ray spectrum from single-energy CT projections using a spectrum library and an implicit neural network, reducing beam-hardening errors compared with two-step image-domain decomposition.","lead":"This paper introduces JSover, a method that estimates both the X-ray energy spectrum and the volume fractions of multiple materials directly from single-energy CT measurements, skipping the usual image reconstruction step. It could make quantitative material decomposition available on conventional CT scanners, where spectral CT hardware is not available.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The joint spectrum/material objective is admitted to be underdetermined; no identifiability analysis shows that the SoftMax-library spectrum is recoverable from single-energy projections, so reported spectrum accuracy may depend on initialization and library choice.","rationale":"The reader's weakest assumption (the ideal-solution basis-material model) is a real physical limitation, but it is shared with all MMD methods, including the TMA and MSC baselines, so it does not uniquely attack the joint-estimation claim. The more load-bearing concern is identifiability: the paper's own text admits that Eq. (9) has numerous feasible solutions, and neither the SoftMax spectrum parameterisation nor the INR prior is accompanied by a uniqueness or stability argument. In the low-attenuation regime the data constrain the spectrum only through a small number of weighted attenuation integrals, so the reported close spectrum matches may reflect library coverage and initialisation rather than data identifiability. I still credit the paper with a physically motivated forward model, large improvements over the two image-domain baselines on simulated and real solution phantoms, and a credible runtime advantage. These strengths support the conditional verdict rather than rejection: the central claim is plausible but unproven until the identifiability experiment is run and error bars over initialisations and noise are reported. I therefore keep the reader's CONDITIONAL verdict, now conditioned on the identifiability check rather than only on the ideal-solution assumption.","tokens_in":15254,"tokens_out":9232,"duration_ms":109789,"concrete_test":"Take the XCAT phantom A and reduce all densities by a factor of 10 so that exp(-A) is approximately 1 - A. Solve for two different spectra w and w' in the same 10-spectrum library with identical weighted-LAC vectors bar-mu_i = sum_k w_k integral eta_k(E) mu_i(E) dE for i=1..4 and sum_k w_k = sum_k w'_k = 1 (a non-trivial null vector of the 5 by 10 constraint matrix). Run JSover with w and w' as different spectrum initialisations. If both runs reach a similarly low L_DC but produce clearly different estimated spectra and materially different material maps, then Eq. (9) is not identifiable from SECT projections and the claimed spectrum accuracy is not guaranteed. Report the two (spectrum, map) pairs and their L_DC values.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III-B explicitly states that Eq. (9) is 'highly underdetermined, with numerous feasible solutions'; the proposed cure is a library-constrained spectrum and an INR prior, with no uniqueness or stability analysis. The danger is concrete. Linearizing Eq. (8) for weakly attenuating objects, exp(-A) is approximately 1 minus A, and the projection depends on the spectrum only through the M weighted LAC integrals bar-mu_i = integral eta(E) mu_i(E) dE. The SoftMax model has N-1 free weights (N=10 in the figures), and low-attenuation data carry roughly M such integral constraints plus weak higher-order beam-hardening terms, so different library spectra with the same bar-mu vector give nearly identical projections. Consequently, the spectrum estimate can be driven by the library prior and initialisation rather than by the data. The reported spectrum MAEs (Table IV, about 0.009) come from simulations in which the same SPEKTR toolkit generates both the library and the GT spectrum, and the only real human-body experiment has no quantitative spectrum ground truth (Section IV-F). The central claim that spectrum and material maps are jointly estimated from SECT projections therefore rests on an unexamined identifiability assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes JSover, a one-step optimization framework for jointly estimating the X-ray energy spectrum and multi-material volume fraction maps directly from single-energy CT projections. The forward model is the standard polychromatic CT model with an ideal-solution mixture assumption; the spectrum is represented as a SoftMax-weighted combination of a precomputed library of spectra, and the material maps are represented by an MLP with hash encoding. The method is optimized with a data-consistency loss using backpropagation. Experiments on simulated XCAT phantoms and on real cone-beam and clinical CT data compare JSover against two image-domain baselines, TMA and MSC, reporting substantially lower RMSE for the decomposed maps and shorter runtime, together with qualitative spectrum estimates that match reference spectra.","tokens_in":15501,"tokens_out":6045,"duration_ms":59425,"significance":"If the central claims hold, JSover is a practically valuable contribution: it replaces the two-step FBP-then-decompose pipeline with a single projection-domain optimization, avoids beam-hardening artifacts by construction, requires no external training data, and reports large accuracy gains on simulated and small real phantoms (RMSE reductions from roughly 0.1 to 0.014–0.027). The forward model is standard physics, and the implementation is transparent and reproducible, with consistent hyperparameters and PyTorch code. The spectrum-estimation component is the least supported part of the contribution: no identifiability analysis is given, and the reported spectrum accuracy is only qualitative in the main experiments; the clinical human-body experiment lacks quantitative validation. These issues are fixable with additional analysis and experiments, so the work is potentially significant but currently not fully substantiated.","major_comments":[{"comment":"The objective in Eq. (9) is acknowledged to be highly underdetermined, but no identifiability or uniqueness analysis is provided for the spectrum-estimation component. For weakly attenuating objects, the linearized forward model depends on the spectrum only through the M weighted LAC integrals ∫η(E) μ_i(E) dE, so multiple library spectra with the same weighted-mean LAC vector produce nearly identical projections; the reported spectrum accuracy in Table IV and Fig. 7 may therefore reflect the library prior and initialization rather than information in the data. Please provide a sensitivity analysis (e.g., varying the initialization of γ and the composition of the library) and report spectrum MAE quantitatively for the main simulated experiments, not only in the architecture ablation.","section":"III-B, Eq. (9)-(11)"},{"comment":"The text states that 'JSover-TV and JSover-INR achieve around 0.03 and 0.02, respectively', but on Phantom B the RMSE of JSover-TV (0.0223) is lower than that of JSover-INR (0.0272). This is the opposite of the stated ordering and weakens the claim that the INR representation enhances decomposition quality; please correct the description or explain the discrepancy (e.g., noise sensitivity of the INR variant).","section":"Table II, Section IV-D"},{"comment":"The real human-body experiment lacks any quantitative ground truth for either the material fractions or the spectrum, so the statement that it 'demonstrates the reliability of our method' is not supported. Moreover, the ideal-solution assumption of Eqs. (3)-(7) is unlikely to hold for mixtures of soft tissue, bone, and air in vivo; the effect of this model mismatch on the estimated fractions and spectrum should be discussed or tested on a numerical phantom with non-ideal mixtures.","section":"Section IV-F"},{"comment":"The reference spectrum for the real solution phantom is a 70 kVP Cu-filtered spectrum, while the library in Fig. 1 consists of 120 kVP and 80 kVP Al-filtered spectra; a convex combination of the library spectra cannot represent the reference exactly. The claimed close match should be quantified (e.g., MAE), and the effect of this representational gap on the spectrum estimate should be discussed.","section":"Section IV-E, Fig. 10"}],"minor_comments":[{"comment":"There are several typos: 'Uncostrained Spectra Esimation' and 'uncostrained' should be 'unconstrained', 'spctra' should be 'spectra', and 'Rerepresentation' should be 'Representation'.","section":"III-B1, III-C, II-C"},{"comment":"The reference to 'Table I' for the quantitative SEMMD results should be 'Table II'.","section":"IV-D"},{"comment":"The motivation for INR is its low-frequency spectral bias, but the best-performing variant uses hash encoding, which is normally associated with high-frequency detail; please reconcile the stated motivation with the architecture choice.","section":"III-B2, IV-G"},{"comment":"Only two image-domain baselines (TMA and MSC) are compared; including a projection-domain or joint-estimation SEMMD baseline would better support the 'state-of-the-art' claim.","section":"IV-C"},{"comment":"The TV regularizer is defined along X-ray coordinates rather than on the image grid; this is unusual and should be justified or replaced by a standard spatial TV prior.","section":"Eq. (16)"},{"comment":"The notation R denotes the full ray set in Eq. (9) but a random subset in Eq. (14); please use distinct symbols to avoid confusion.","section":"Eq. (9) and Eq. (14)"},{"comment":"Table II reports RMSE without standard deviations, while Tables IV and V report mean±std; please clarify how many repetitions are performed and whether the RMSE values are averaged over the two phantoms.","section":"Table II vs Tables IV/V"}],"recommendation":"major_revision","confidential_remarks":"The identifiability issue is the main risk; I would not reject because the material-decomposition gains over the two baselines are large and the method is reproducible, but the spectrum-estimation claim needs stronger support. The comparison to only two baselines and the lack of quantitative metrics in the human-body experiment also limit the paper's current strength. The Table II discrepancy between JSover-TV and JSover-INR on Phantom B should be corrected before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: JSover is the first one-step projection-domain SEMMD method that estimates the spectrum and material fractions jointly using an INR, and on its own experiments it beats the image-domain optimization baselines by a large margin. That is real. The forward model is standard polychromatic CT, the ideal-solution basis expansion is clearly derived, and the SoftMax reparameterization turns a constrained spectrum search into an unconstrained one without extra regularizers. The INR solver also helps stabilize volume fractions. The reported RMSE drop from ~0.1 (TMA/MSC) to ~0.014–0.027 on simulated XCAT phantoms is substantial, and the real solution-phantom results show decomposition accuracy that the baselines do not reach.\n\nThe soft spot, and it is central to one of the two advertised outputs, is spectrum identifiability. The paper itself says Eq. (9) is highly underdetermined with numerous feasible solutions, then answers with the spectrum library and the INR prior. That constrains the solution but does not establish that the data pick the true spectrum. For weakly attenuating objects the projection depends on the spectrum mainly through energy-averaged LAC integrals, so many library spectra can produce nearly identical projections. The reported spectrum MAEs (~0.009) come from simulations where the same SPEKTR toolkit generates both the library and the ground truth, and the real human-body experiment has no spectrum ground truth at all. The real solution phantom compares against a SPEKTR reference based on known settings, not an independent measurement. So I would not treat the spectrum estimates as demonstrated fact; I would treat them as plausible and library-dependent. This does not sink the decomposition claim, which is the more useful part, but the paper should either add an identifiability/stability analysis or soften the joint-estimation claim.\n\nTwo smaller points. The comparisons are only against optimization-based SEMMD (TMA, MSC) and an ablated TV variant; supervised DL baselines are named in the related work but not evaluated, so \"state-of-the-art\" is too strong. And there is no code or data release, which matters more than usual here because spectrum reproducibility depends on the exact SPEKTR libraries and preprocessing; the implementation details are mostly there but not all.\n\nOverall: solid within its scope, with one load-bearing gap in the spectrum story and one overreach in the baseline claim. It deserves review, and the authors should be asked for identifiability analysis, error bars over multiple runs, a supervised baseline, and code/data before acceptance.","headline":"A clever one-step single-energy MMD framework with strong empirical gains, but the joint spectrum estimator lacks an identifiability analysis and the spectrum validation is partly circular.","tokens_in":16018,"tokens_out":2965,"would_cite":true,"duration_ms":30096,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Jointly estimating the X-ray spectrum and material volume fractions directly from single-energy CT projections outperforms two-step SEMMD methods in accuracy and speed.","keywords":["multi-material decomposition","single-energy CT","spectrum estimation","implicit neural representation","beam hardening","unsupervised learning","CT reconstruction"],"falsifier":"A phantom containing a known material outside the predefined basis set (e.g., a two-material mixture plus a trace of a third element) would test the claim: if JSover still fits the projections well but the recovered volume fractions are systematically biased, the ideal-solution assumption is violated and the joint estimate returns a plausible but wrong decomposition.","tokens_in":15065,"feed_emoji":"🩻","tokens_out":4596,"duration_ms":41651,"temperature":0.7,"pith_summary":"This paper tackles a central limitation of conventional CT: a single-energy scan cannot easily reveal what materials are inside the body because the X-ray spectrum is unknown and attenuation depends on energy. The authors propose JSover, a one-step optimization that jointly recovers the X-ray energy spectrum and the volume fractions of several predefined materials directly from raw projections. Unlike the standard two-step pipeline—first reconstruct a monochromatic image, then decompose—this approach builds polychromatic beam-hardening physics into the forward operator, so decomposition happens before reconstruction artifacts are introduced. On simulated and real phantoms, the method achieves an order-of-magnitude improvement in material fraction accuracy while also estimating the spectrum, in a fraction of the runtime of existing two-step methods.","feed_headline":"One step turns single-energy CT scans into material maps","feed_subtitle":"Jointly estimating the X-ray spectrum and tissue fractions from raw projections beats two-step decomposition.","key_machinery":"The load-bearing object is the polychromatic forward model $\\widetilde{H}$ of Eq. (13), which generates a simulated projection $\\hat{\\rho}(r)$ by integrating over discrete energies the exponential of a sum of basis-material line integrals weighted by a softmax-combined spectrum $\\eta(E)=\\sum_i \\mathrm{SoftMax}(\\gamma_i)\\eta_i(E)$. This single differentiable operator ties the two unknowns together: an MLP $F_\\Phi:\\mathbb{R}^3\\to\\mathbb{R}^M$ maps spatial coordinates to volume fractions $\\alpha(x)$, and the spectrum parameters $\\gamma$ are optimized jointly by backpropagation through the same loss. The softmax transformation guarantees non-negativity and unit sum without constraints, and the hash-encoded MLP imposes a low-frequency inductive bias that regularizes the ill-posed inverse problem.","core_discovery":"The central claim is that single-energy CT projections contain enough information to estimate both the X-ray spectrum and the volume fractions of M known basis materials, provided the two unknowns are solved jointly under a physics-consistent polychromatic forward model. The paper formalizes this as minimizing a data-consistency loss between measured and simulated projections, with the spectrum represented as a softmax-weighted combination of library spectra (making the optimization unconstrained) and the material maps represented by an implicit neural network. On simulated XCAT phantoms the method reaches volume-fraction RMSE of 0.014–0.027, versus roughly 0.1 for two-step baselines TMA and MSC, and it recovers the reference spectrum closely; on real solution phantoms and a clinical human-body phantom it produces anatomically plausible decompositions. The authors state this is the first unsupervised deep-learning approach to joint spectrum estimation and single-energy multi-material decomposition.","pith_inferences":["If the basis-material set were expanded or made learnable, JSover's framework could become a general material-mapping tool beyond the predefined adipose, muscle, bone, and air set.","The softmax spectrum representation might be reused for beam-hardening correction in standard CT reconstruction, since it estimates the spectrum without any calibration step.","A straightforward test extension would be to run JSover on a phantom with a known three-material mixture and check whether the recovered volume fractions are unbiased; the paper's real phantom experiments only used two-component solutions.","The INR solver's low-frequency bias, while helpful for regularization, may limit spatial resolution in fine structures; this could be tested by comparing against the TV-regularized variant on a resolution phantom."],"forward_implications":["Material decomposition becomes possible on existing single-energy scanners without pre-measured spectra, removing a major barrier to clinical adoption.","Because the forward model is polychromatic, the decomposition is performed before any FBP reconstruction artifacts enter the process, eliminating beam-hardening artifacts at the source.","The spectrum estimate is a free by-product of the same optimization, useful for scanner calibration and dose modeling.","The method degrades gracefully under undersampled projections (down to 4x fewer views), so it could reduce radiation dose.","The unsupervised nature means no paired spectral-CT training data is required, so it can transfer to new scanners or body regions without retraining."],"supporting_citations":[{"why":"The TMA baseline: an optimization-based two-step SEMMD method whose accuracy JSover is compared against.","marker":"[8]"},{"why":"The MSC baseline: a two-step SEMMD method adding material sparsity constraints, used as the other representative comparison.","marker":"[7]"},{"why":"Establishes the ideal-solution assumption that a mixture's mass attenuation coefficient is a volume-fraction-weighted sum of basis material MACs.","marker":"[3]"},{"why":"A library-based spectrum estimation method whose weighted-sum spectrum model JSover extends with a softmax transformation.","marker":"[14]"},{"why":"Documents the spectral bias of neural networks toward low-frequency functions, the inductive prior that regularizes the INR material maps.","marker":"[18]"},{"why":"Provides the hash encoding architecture used in the MLP for efficient coordinate-based mapping.","marker":"[37]"},{"why":"Generates the spectrum libraries and reference spectra used in both simulated and real experiments.","marker":"[36]"},{"why":"FBP, the standard reconstruction algorithm whose monochromatic output is the input to two-step SEMMD baselines.","marker":"[13]"}],"fun_headline_variants":["One step, both unknowns: spectrum and materials from single-energy CT","One-step joint estimation from single-energy CT beats two-step","Implicit neural network enables one-step single-energy CT decomposition","Unsupervised one-step SEMMD: joint spectrum and material maps","Single-energy CT simulates spectral decomposition in one step"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forward model assumes the scanned object is an ideal mixture of M predefined basis materials with known energy-dependent attenuation, so that the mixture's attenuation is a volume-fraction-weighted sum; real tissues are not always ideal mixtures, and the clinical experiment assumes only soft tissue, bone, and air.","fun_headline_variants_meta":{"raw":{"variants":["One step, both unknowns: spectrum and materials from single-energy CT","One-step joint estimation from single-energy CT beats two-step","Implicit neural network enables one-step single-energy CT decomposition","Unsupervised one-step SEMMD: joint spectrum and material maps","Single-energy CT simulates spectral decomposition in one step"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0012,"raw_usage":{"total_tokens":4973,"prompt_tokens":1001,"completion_tokens":3972,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":3887}},"tokens_in":617,"tokens_out":3972,"duration_ms":27553,"temperature":1.0,"reasoning_tokens":3887,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:03:02.183182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A phantom containing a known material outside the predefined basis set (e.g., a two-material mixture plus a trace of a third element) would test the claim: if JSover still fits the projections well but the recovered volume fractions are systematically biased, the ideal-solution assumption is violated and the joint estimate returns a plausible but wrong decomposition.","supporting_citations":[{"cited_title":"Image domain multi-material decomposition using single energy ct,","cited_arxiv_id":null,"evidence_quote":"The TMA baseline: an optimization-based two-step SEMMD method whose accuracy JSover is compared against."},{"cited_title":"Multi-material decomposition for single energy ct using material sparsity constraint,","cited_arxiv_id":null,"evidence_quote":"The MSC baseline: a two-step SEMMD method adding material sparsity constraints, used as the other representative comparison."},{"cited_title":"A flexible method for multi-material decomposition of dual-energy ct images,","cited_arxiv_id":null,"evidence_quote":"Establishes the ideal-solution assumption that a mixture's mass attenuation coefficient is a volume-fraction-weighted sum of basis material MACs."},{"cited_title":"An indirect transmission measurement-based spectrum estimation method for computed tomog- raphy,","cited_arxiv_id":null,"evidence_quote":"A library-based spectrum estimation method whose weighted-sum spectrum model JSover extends with a softmax transformation."},{"cited_title":"On the spectral bias of neural networks,","cited_arxiv_id":null,"evidence_quote":"Documents the spectral bias of neural networks toward low-frequency functions, the inductive prior that regularizes the INR material maps."},{"cited_title":"Technical note: spektr 3.0—a computational tool for x-ray spectrum,","cited_arxiv_id":null,"evidence_quote":"Generates the spectrum libraries and reference spectra used in both simulated and real experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"FBP, the standard reconstruction algorithm whose monochromatic output is the input to two-step SEMMD baselines."}],"review_version":1}