{"id":"d054330c-e477-4b33-9338-935cfc97c702","arxiv_id":"2507.04665","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Adding an STFT-based spectral loss to a conditional GAN improves high-frequency fidelity of generated machining force signals and, used as training augmentation, cuts surface roughness prediction MAPE from 31.4% to 8.8%.","lead":"This paper combines a conditional GAN with a spectral (frequency-domain) loss to create synthetic cutting-force signals for ultra-precision machining, then uses those signals to expand a 52-sample training set for surface roughness prediction. The authors report that adding 520 synthetic samples lowers prediction error from about 31% to about 9%.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline MAPE gain depends on an unstated pairing of synthetic signals with machining parameters; the generator in Eqs. 1-3 is conditioned only on noise and Ra label, so the augmented signal-parameter-roughness triples are not defined.","rationale":"The reader's weakest assumption identifies the same load-bearing gap: the paper's equations for the generator omit the machining parameters that the prediction phase treats as inputs. This is not a matter of disagreeing with current consensus; it is an internal inconsistency in the data-generation protocol. The strongest claim would fail if the pairing is arbitrary, and no part of the manuscript supplies the missing protocol. Additional evidence supports caution: the train-set size is reported as both 52 and 56 samples (Section 3.1 vs. Section 3.3), the test set has only 12 samples, and no code, data, error bars, or repeated splits are provided. The wavelet-coherence metric only measures similarity to training-set signals and does not validate whether the generated waveforms are physically associated with the attached parameters. Because the central result cannot be checked as stated and the load-bearing pairing assumption is unstated, maintaining the reader's REJECT verdict is appropriate; the proposed shuffle ablation would settle whether the pairing assumption actually matters.","tokens_in":13553,"tokens_out":5457,"duration_ms":58564,"concrete_test":"Re-run the Section 3.3 experiment with the parameter columns (spindle speed, feed rate, depth of cut) of the 520 synthetic samples randomly permuted before training, keeping waveforms and Ra labels fixed; train the CNN-Transformer with the same protocol and report MAPE on the same 12-sample test set. If the shuffled-parameter MAPE remains near 8.8%, the reported improvement does not depend on the claimed signal-parameter pairing and the central claim is unsupported. If MAPE degrades substantially, request the exact pairing rule and repeat the comparison using that rule, then re-evaluate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section 3.3: MAPE from 31.4% to 8.8% for CNN-Transformer after adding 520 generated samples) presupposes that every augmented training item is a valid triple (force signal, machining parameters, Ra). But the HAS-CGAN generator in Section 2.2.1 (Eqs. 1-3) takes as input only sinusoidal latent noise and a surface-roughness label; spindle speed, feed rate, and depth of cut do not enter the generation process. Section 2.1 nonetheless states that augmented data contain generated signals 'with corresponding machining parameters and labels,' and Section 3.3 reports training on those triples. The paper never specifies how parameters are assigned to synthetic signals. If the parameters are assigned arbitrarily, or matched only by Ra label, then the augmented training distribution contains input-output combinations that are not physically coherent: the same synthetic waveform can be paired with multiple parameter sets while the label is fixed by construction. The 8.8% MAPE could then reflect the predictor exploiting label/artifact correlations in the generated set rather than learning the signal-parameter-roughness mapping that the paper claims to improve. This is load-bearing because, without valid triples, the before/after comparison does not establish that data augmentation works; it only shows that a different, possibly corrupted, training distribution achieves lower error on a 12-sample test set.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HAS-CGAN, a conditional GAN augmented with a spectral loss, to generate synthetic 1D cutting-force signals for data augmentation in ultra-precision machining surface roughness prediction. The authors compare five CGAN variants on signal fidelity using wavelet coherence, then combine generated signals with machining parameters to enlarge a 52-sample training set and report that training CNN-Transformer on a 10x-augmented dataset reduces MAPE from 31.4% to 8.8%. The central claim is that the generated signals are realistic enough to serve as effective training data.","tokens_in":13852,"tokens_out":7990,"duration_ms":86372,"significance":"If substantiated, the approach would be practically valuable for small-sample industrial datasets, where collecting and labeling sensor data is expensive. The paper has some strengths: it systematically compares five GAN variants, uses a quantitative fidelity metric (wavelet coherence), and evaluates augmentation across multiple prediction architectures. However, the central claim is not supported as written. The generator is not conditioned on machining parameters, yet the augmented dataset pairs each synthetic signal with machining parameters; this gap undermines the validity of the reported MAPE improvement. Internal inconsistencies in dataset counts and augmentation sizes, the circularity of the wavelet-coherence fidelity claim, and the absence of any uncertainty quantification further reduce confidence. The manuscript is not ready for publication in its current form.","major_comments":[{"comment":"The generator takes only sinusoidal noise and a surface-roughness label as input; spindle speed, feed rate, and depth of cut never enter the generation process. Yet §2.1 states that generated signals are provided \"with corresponding machining parameters,\" and §3.3 trains predictors on (signal, parameters, Ra) triples. The paper never specifies how machining parameters are assigned to synthetic signals. If the assignment is arbitrary or based only on the Ra label, the augmented triples are not physically coherent because the generated waveform has no parameter dependence. The reported MAPE improvement from 31.4% to 8.8% could then reflect label/artifact correlations rather than a learned signal-parameter-roughness mapping. This is load-bearing because the central claim that data augmentation improves prediction is meaningless without valid triples.","section":"§2.2.1 (Eqs. 1–6) and §3.3"},{"comment":"The training-set size is stated as 52 samples in §3.1, but §3.3 refers to \"the original 56 real samples.\" Similarly, the 5-times augmentation is described as \"210 samples\" in the text but as \"260 samples\" in the Fig. 3.c caption. Since 10x = 520 is consistent with 52 samples, the 5x and 56-sample counts are internally contradictory. These inconsistencies make the exact experimental protocol ambiguous and undermine the quantitative precision of the headline result.","section":"§3.1 and §3.3"},{"comment":"Wavelet coherence is used as the principal fidelity metric, with conflicting thresholds reported in the abstract (>0.85), conclusion (>0.9), and text (>0.8). More importantly, the high WC for HAS-CGAN is partly by construction: the generator's custom loss in Eq. (5) directly minimizes the magnitude difference between the STFT of real and generated signals, and wavelet coherence measures time-frequency similarity. Without a baseline comparison (e.g., WC between two real signals, or WC for a non-spectral-loss CGAN), the fidelity gain could be an artifact of explicit frequency-domain matching rather than genuine signal realism. The downstream MAPE reduction is an independent outcome, but the generation-fidelity claim is not established.","section":"§3.2 and Eq. (5)"},{"comment":"The evaluation relies on a single random 52/12 split with no confidence intervals, no repeated runs, and no code. The improvement in MAPE from 31.4% to 8.8% is based on a single 12-sample test set; the paper does not report variance across seeds or any statistical significance test. Without this, the improvement is anecdotal rather than established, especially given the very small test set.","section":"§3.3"}],"minor_comments":[{"comment":"The paper refers to \"HAS-CGAN (Hierarchical Attention-Supervised Conditional Generative Adversarial Network)\", which contradicts the title's \"Hybrid Adversarial Spectral Loss\"; please use one consistent name.","section":"§3.1"},{"comment":"The STFT subscripts/superscripts for the real and generated signals are difficult to distinguish; both appear as x_i with different labels, and the notation should be clarified.","section":"Eq. (5)"},{"comment":"The generator is described as using \"three 1D-convolutional transpose layers\" in §2.1 but as a \"3-layer fully connected network\" in §3.2; please reconcile this discrepancy.","section":"§2.1 vs §3.2"},{"comment":"The order of methods is inconsistent: \"SVR, LSTM and RF\" appears in one paragraph and \"SVR, RF and LSTM\" in the next; keep the ordering consistent throughout.","section":"§3.3"},{"comment":"The wavelet-coherence claims differ across the abstract (>0.85), conclusion (>0.9), and body (>0.8); report the actual values or a range associated with the relevant figure.","section":"Abstract/Conclusions"},{"comment":"There are numerous typographical and grammatical errors (e.g., \"no matther,\" \"augemented,\" \"resepectively\") that interfere with readability; a thorough language edit is needed.","section":"General"}],"recommendation":"reject","confidential_remarks":"The core methodological gap — the generator ignores machining parameters even though the augmented triples include them — is not a minor omission but a fundamental flaw. Fixing it would require redesigning the method and rerunning all experiments, which goes beyond a normal revision. The internal numeric inconsistencies and lack of any uncertainty quantification further weaken confidence in the reported results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper does something useful for the UPM community: it compares five CGAN variants on 1D force signals and shows that a simple convolutional CGAN with an added STFT magnitude loss produces higher wavelet coherence, especially for high-frequency signals. That part is believable and the figures support it. Second, the headline augmentation result—MAPE down from 31.4% to 8.8%—is not established. The generator conditions only on noise and an Ra label, but the augmented training set is said to contain 'corresponding machining parameters.' The paper never says where those parameters come from or how the signal-parameter-roughness triples are formed. If the pairing is arbitrary, the augmentation corrupts the input-output mapping, and the reported gain is meaningless. This is load-bearing.\n\nWhat's genuinely good: the systematic comparison of CGAN/ACGAN/WCGAN/convolutional CGAN on a real manufacturing dataset, with a sensible argument that complex adversarial variants overfit small 1D datasets. The spectral-loss trick is standard in audio GANs, and the paper does not cite that line of work, so the novelty claim is overstated. But the application and the empirical comparison have value.\n\nSoft spots, in order of severity. (1) The missing pairing protocol. (2) Internal contradictions: the abstract says 52 samples, Section 3.1 says 52 training and 12 test, but Section 3.3 says 56 real samples; 5x augmentation is listed as 210 samples instead of 260; wavelet coherence is reported as >0.85 in the abstract and >0.9 in the conclusion. These are small but they erode trust. (3) No code, no data, no confidence intervals, a single 52/12 split. For a claim whose whole point is statistical improvement, that is thin.\n\nThe paper is not a waste of time. The comparative result and the diagnostic that spectral loss helps high-frequency fidelity are worth an editor's attention. But the central augmentation claim needs a clearly specified pairing rule and a more rigorous evaluation before it can be accepted.\n\nRecommendation: send to peer review, but expect heavy revision. The niche is real, the flaws are fixable, and a good reviewer can turn this into a solid empirical paper. I would not cite it in its current form.","headline":"Useful empirical comparison of CGAN variants for 1D force signals, but the headline augmentation claim rests on an unspecified signal-to-parameter pairing and a single small split.","tokens_in":14372,"tokens_out":2270,"would_cite":false,"duration_ms":22467,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that augmenting a 52-sample ultra-precision machining dataset with 520 generated force signals lowers surface roughness prediction error from 31.4% to about 8.8% MAPE.","keywords":["signal data augmentation","conditional generative adversarial network","spectral loss","surface roughness prediction","ultra-precision machining","wavelet coherence","CNN-Transformer"],"falsifier":"A concrete check would be to train the same CNN-Transformer predictor after randomly permuting the machining parameters attached to the generated signals while keeping the roughness labels fixed; if the MAPE stays near 8.8%, the improvement does not depend on an intact signal-parameter relationship and the central claim is weakened.","tokens_in":13356,"feed_emoji":"⚙️","tokens_out":4523,"duration_ms":49294,"temperature":0.7,"pith_summary":"The paper tries to establish that a conditional generative adversarial network with an added frequency-domain spectral loss can generate realistic ultra-precision machining force signals, and that those synthetic signals can serve as effective extra training data for surface roughness prediction. On a 64-sample milling dataset, the authors compare five CGAN variants and report that their hybrid adversarial spectral loss reproduces high-frequency content best, with wavelet coherence above 0.85. Adding ten times as many generated samples lowers the CNN-Transformer mean absolute percentage error from 31.4% to about 8.8%, with diminishing returns beyond that. This matters because real ultra-precision machining data is scarce and expensive to collect; if synthetic force signals carry usable physical information, accurate quality-control models could be trained without extensive additional machining experiments.","feed_headline":"520 synthetic signals cut roughness error from 31.4% to 8.8%","feed_subtitle":"A conditional GAN with spectral loss supplies training data for ultra-precision machining models without extra cutting.","key_machinery":"The central object is the hybrid adversarial spectral loss for the generator, defined as $HAS\\_Loss_G = \\gamma_1 Loss_1 + \\gamma_2 Loss_2$ with $\\gamma_1 + \\gamma_2 = 1$, where $Loss_1$ is the standard conditional GAN generator loss and $Loss_2$ averages, over the batch and time frames, the squared Frobenius norm of the difference between the STFT magnitudes of real and generated force signals. This spectral term constrains the generator in the frequency domain so that high-frequency content is not lost. The generator itself is a three-layer 1D convolution-transpose network that takes sinusoidal noise combined with a surface-roughness label as input, while the discriminator is a three-layer 1D convolutional network that judges whether a signal is real or generated under the same label condition.","core_discovery":"The paper's central claim is that a lightweight 1D convolutional conditional GAN, trained with a generator loss that penalizes the difference between the short-time Fourier transform magnitudes of real and generated signals, produces synthetic force signals that are more faithful to real machining signals than those from plain CGAN, convolutional CGAN, ACGAN, or WCGAN. The improvement is most visible for high-frequency signals, where the spectral loss directly punishes fitting error in the Fourier domain. When these generated signals are combined with machining parameters and added to the training set, end-to-end predictors that extract features automatically improve substantially, with the best model reaching about 8.8% MAPE compared with 31.4% on the original 52 real samples. The paper concludes that for 1D industrial signals, simpler GAN architectures with frequency-aware losses are more suitable than theoretically more complex variants, and that CGAN-based augmentation is a viable route toward real-time virtual metrology for ultra-precision machining.","pith_inferences":["Editorial extension: because the generator conditions only on sinusoidal noise and a roughness label, the paper leaves open how each synthetic signal is paired with spindle speed, feed rate, and depth of cut; conditioning the generator on those parameters explicitly would make the augmentation mechanism more transparent and testable.","Editorial extension: part of the reported improvement could come from the predictor seeing more examples spread across the label range rather than from physically faithful waveforms; a control experiment using randomly relabeled or noise-only augmented data would separate these effects.","Editorial extension: the same spectral-loss recipe may transfer to other small-dataset industrial signal problems, but the optimal augmentation ratio is likely to depend on signal dimensionality and label diversity rather than being a universal tenfold rule."],"forward_implications":["End-to-end models that learn features from raw waveforms benefit from generated signals, while models that rely on hand-crafted time- and frequency-domain features do not; augmentation helps only when the predictor can use the raw signal structure.","A roughly tenfold augmentation, or about 520 generated samples, is the practical ceiling for this dataset; beyond it the prediction error plateaus near 9% MAPE, so extra generation yields diminishing returns.","The spectral loss is what improves high-frequency fidelity; the paper reports that CGAN variants without it fail to reproduce the amplitudes of high-frequency force signals.","If the result holds, CGAN-based augmentation offers a route toward virtual metrology for ultra-precision machining, reducing dependence on time-consuming offline surface measurements."],"supporting_citations":[{"why":"Supplies the 64-sample ultra-precision milling dataset, including force signals, machining parameters, and measured Ra labels, that all experiments use.","marker":"[14]"},{"why":"Provides the original GAN adversarial training formulation and the generator/discriminator structure that the proposed network adapts.","marker":"[26]"},{"why":"Introduces conditional GANs, the label-conditioned training scheme that lets generated signals carry surface roughness labels.","marker":"[42]"},{"why":"Provides the ACGAN baseline with an auxiliary classifier loss that the paper compares against.","marker":"[35]"},{"why":"Provides the improved Wasserstein GAN with gradient penalty that underlies the WCGAN baseline and its comparison.","marker":"[36]"},{"why":"Establishes the preprocessing procedures used for the traditional feature-based prediction models in the comparison.","marker":"[41]"}],"fun_headline_variants":["Spectral-loss GAN cuts machining roughness error to 9%","Frequency-aware GAN shrinks machining prediction error to 9%","Synthetic force signals via spectral CGAN reduce roughness error to 9%","Spectral-loss CGAN beats five rivals in machining data augmentation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each synthetic force signal can be paired with machining parameters such as spindle speed, feed rate, and depth of cut in a way that preserves the true signal-parameter-roughness relationship, but the paper never states where those parameter values come from or how the pairing is made.","fun_headline_variants_meta":{"raw":{"variants":["Spectral-loss GAN cuts machining roughness error to 9%","Frequency-aware GAN shrinks machining prediction error to 9%","Synthetic force signals via spectral CGAN reduce roughness error to 9%","Spectral-loss CGAN beats five rivals in machining data augmentation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000562,"raw_usage":{"total_tokens":2647,"prompt_tokens":902,"completion_tokens":1745,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":1669}},"tokens_in":518,"tokens_out":1745,"duration_ms":13724,"temperature":1.0,"reasoning_tokens":1669,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:42:49.226450+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check would be to train the same CNN-Transformer predictor after randomly permuting the machining parameters attached to the generated signals while keeping the roughness labels fixed; if the MAPE stays near 8.8%, the improvement does not depend on an intact signal-parameter relationship and the central claim is weakened.","supporting_citations":[{"cited_title":"F., & Zheng, P","cited_arxiv_id":null,"evidence_quote":"Supplies the 64-sample ultra-precision milling dataset, including force signals, machining parameters, and measured Ra labels, that all experiments use."},{"cited_title":"& Bengio, Y","cited_arxiv_id":null,"evidence_quote":"Provides the original GAN adversarial training formulation and the generator/discriminator structure that the proposed network adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ACGAN baseline with an auxiliary classifier loss that the paper compares against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the improved Wasserstein GAN with gradient penalty that underlies the WCGAN baseline and its comparison."},{"cited_title":"(2024) Roughness prediction of end milling surface for behavior mapping of digital twined machine tools [version 2; peer review: 2 approved, 1 approved with reservations]","cited_arxiv_id":null,"evidence_quote":"Establishes the preprocessing procedures used for the traditional feature-based prediction models in the comparison."}],"review_version":1}