{"id":"2c379d0f-98e4-4567-a3a0-4ca61f1b955e","arxiv_id":"2608.07363","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"QFCQT claims up to 43.9% MSE improvement over HAT on ETTh2 via Lee-oscillator gated activation, but internal table inconsistencies and undefined ablation components undermine the claim.","lead":"QFCQT adds a chaotic, oscillator-based activation to a transformer to predict volatile series like power temperatures and stock prices, reporting large MSE wins over baselines. The paper's own tables contain duplicate numbers and negative improvements that contradict the 'consistent outperformance' headline.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own Table IV contradicts the central claim of consistent outperformance: QFCQT is worse than COTN on ETTh1 h48/h336 MAE and ETTh2 h336 MSE, and Table II duplicates the QFCQT rows for ETTh1 h48 and h168.","rationale":"I focused on the internal contradiction in the reported results rather than the matched-capacity confound because it is more load-bearing. The central claim is empirical superiority; a single negative result in the paper's own comparison table defeats the 'consistently' qualifier, and the duplicated ETTh1 rows call the integrity of the reported numbers into question. This concern is decisive regardless of parameter counts or mechanistic attribution: even a perfectly capacity-matched study would still need to explain why the abstract says 'consistently' when Table IV shows otherwise. The reader's stated weakest_assumption was the absence of a matched-capacity control, but their rationale also pointed to Table IV contradictions; I therefore mark agreement as partial. My proposed check is a direct recomputation plus targeted rerun that would settle whether the negative entries and duplicated rows are genuine or artifacts of transcription. If they are genuine, the central claim must be revised; if they are artifacts, the manuscript still needs correction and release of code and seeds before the central claim can be accepted. The reader's REJECT verdict remains appropriate.","tokens_in":9130,"tokens_out":4210,"duration_ms":38410,"concrete_test":"Recompute every relative-improvement entry in Table IV from the raw MSE/MAE values in Table III using (baseline - QFCQT)/baseline, and rerun the ETTh1 h48 and h168 QFCQT configurations with fixed seeds and identical hyperparameters. If the negative relative improvements persist and the two horizon rows remain numerically identical, the manuscript's own data contradict 'consistently outperforms', independent of any capacity or mechanism arguments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and conclusion claim QFCQT 'consistently outperforms' HAT, COTN, and other baselines, and Section V attributes this to chaos-aware activation. However, the paper's own results falsify the 'consistent' claim. Table IV reports negative relative improvements over COTN in three settings: ETTh1 h48 MAE -3.2%, ETTh1 h336 MAE -3.1%, and ETTh2 h336 MSE -1.1%. These are not rounding artifacts: in Table III, QFCQT's MAE is 0.584 versus COTN's 0.566 at ETTh1 h48, and QFCQT's MSE is 2.542 versus COTN's 2.514 at ETTh2 h336. Since 'consistently outperforms' is an all-quantifier claim, any negative setting defeats it as written. Compounding this, Table II lists identical QFCQT results (MSE 0.619, MAE 0.584) for ETTh1 horizons 48 and 168; such duplicated rows indicate a data-handling or reporting error. Without corrected tables, the central empirical claim is unsupported by the manuscript itself. The matched-capacity concern raised by the reader is real but secondary: it would weaken the mechanistic interpretation, whereas this internal contradiction attacks the headline result directly.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes QFCQT, a Transformer forecasting architecture whose feed-forward blocks replace ordinary pointwise activations with a learnable mixture of eight Lee-oscillator activation families, fused with GELU through a gating scalar. The model is evaluated on ETTh1, ETTh2, and a private A-share stock dataset, with the abstract and conclusion claiming that QFCQT 'consistently outperforms' Informer, LogTrans, LSTMa, HAT, COTN, and TimesNet, including large gains at ETTh2 horizon 24. The paper explicitly disclaims any formal quantum or fractal derivation, framing those terms as computational analogies.","tokens_in":9488,"tokens_out":3454,"duration_ms":31119,"significance":"If the empirical claims could be trusted, the core idea—replacing standard static activations in Transformer feed-forward blocks with a learnable, oscillator-based chaotic gate—would be a simple and potentially useful architectural modification for volatile time series. The paper also deserves credit for clearly stating the limits of its quantum/fractal terminology and for providing an ablation that shows some degradation when the proposed modules are removed. However, the headline claim of consistent outperformance is directly contradicted by the paper's own tables, and the mechanistic interpretation is under-supported by the experiments as reported. No code or data are provided, which further limits verification.","major_comments":[{"comment":"The abstract and conclusion state that QFCQT 'consistently outperforms' strong baselines, but Table IV reports negative relative improvements over COTN in three settings: ETTh1 h48 MAE (-3.2%), ETTh1 h336 MAE (-3.1%), and ETTh2 h336 MSE (-1.1%). Table III confirms the underlying comparisons (QFCQT MAE 0.584 vs COTN MAE 0.566 at ETTh1 h48; QFCQT MSE 2.542 vs COTN MSE 2.514 at ETTh2 h336). Since 'consistently' is an all-quantifier claim, the paper's own results falsify it. The claim must be qualified to 'in most settings' or similar, and the negative cases need explicit discussion.","section":"Abstract, Section V.A, Section VII, Tables III and IV"},{"comment":"The QFCQT row for ETTh1 horizon 48 (MSE 0.619, MAE 0.584) is identical to the QFCQT row for ETTh1 horizon 168 (MSE 0.619, MAE 0.584), and Table III repeats the same values for horizon 48. Identical results across two different horizons are implausible and indicate a data-handling or reporting error. These tables must be corrected and the source of the duplication explained before the empirical results can be assessed.","section":"Table II"},{"comment":"The ablation removes QuantumSuperpositionLORS, VectorizedLeeOscillator, FractalModulatedAttention, and the original forecasting head simultaneously, and results are reported only for horizon 24. This design cannot support the statement in Section V.C that 'each of them contributes positively' to forecasting performance; individual ablations and additional horizons are needed. Furthermore, FractalModulatedAttention is not defined in the model description (Equations 8–10 describe only standard multi-head self-attention), so the reader cannot determine what component was actually removed.","section":"Section V.C, Table V"},{"comment":"The mechanistic interpretation in Section V.D attributes the gains to chaos-aware activation, but the model adds eight oscillator families, a learnable mixture, and a gate on top of a baseline without a matched-capacity control. The extra parameters alone could explain the observed improvements. Additionally, the hand-set Lee oscillator coefficients in Table I are chosen without sensitivity analysis, so the assumption that one fixed set of coefficients is universally appropriate across all datasets and horizons is unverified and load-bearing for the claimed mechanism.","section":"Section V.D, Table I"}],"minor_comments":[{"comment":"The title contains a typo: 'V olatile' should be 'Volatile'.","section":"Title"},{"comment":"Reference [7], cited for LogTrans, points to a biomedical image segmentation paper by Nie et al. (2022), not to the LogTrans time-series forecasting work; the correct source should be cited.","section":"References, [7]"},{"comment":"The sentence 'Table II and III reports the comprehensive forecasting results' has a subject-verb agreement error and inconsistent table numbering; it should read 'Tables II and III report...'.","section":"Section V.A"},{"comment":"Equation (4) and Equation (7) both define H(0) = X W_e + b_e; the duplicate definition should be removed or replaced with a cross-reference.","section":"Equations 4 and 7"},{"comment":"The text says 'we use two representative time-series datasets' but then lists three: ETTh1, ETTh2, and A-share Stock Index. This should be corrected.","section":"Section IV.A"}],"recommendation":"major_revision","confidential_remarks":"The paper's heavy reliance on self-cited prior work is not by itself disqualifying, but the duplicated rows in Table II and the contradictions between Table IV and the abstract/conclusion suggest that a careful data audit is needed before the results can be taken at face value. The proposed method is an incremental extension of COTN into a Quantformer-style backbone, and the current evidence does not support the strong 'consistent outperformance' framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is an incremental extension of the group's own COTN line—soft mixture over eight Lee oscillator families inside a Quantformer-style encoder—and the authors are upfront that it's a practical integration, not a new theory. That part is fine. The problem is that the abstract and conclusion claim QFCQT 'consistently outperforms' HAT and COTN, and Table IV in the same paper shows negative relative improvements over COTN on ETTh1 h48 MAE (-3.2%), ETTh1 h336 MAE (-3.1%), and ETTh2 h336 MSE (-1.1%). You can't call that consistent. The stress-test note is correct, and it's not a rounding artifact: the raw numbers in Table III back it up.\n\nThere's also a clear copy-paste error in Table II: the QFCQT row for ETTh1 h48 (MSE 0.619, MAE 0.584) is identical to the h168 row. Same values for two different horizons is a data-handling red flag. The ablation study in Section V.C names modules—QuantumSuperposition LORS, VectorizedLeeOscillator, FractalModulatedAttention—that never appear in the methodology. That mismatch suggests the ablation description was written for a different version of the model. No code, seeds, or dataset splits are provided, so none of the numbers can be checked.\n\nWhat's genuinely new is the learnable soft mixture over eight hand-set oscillator families (Eq. 16–17). That is a reasonable extension of COTN's single-oscillator approach, and the paper does a decent job explaining the mechanics. The explicit disclaimer that 'quantum-fractal-inspired' is just an analogy is a point in its favor. The comparison set includes external baselines like Informer, LogTrans, and TimesNet, so it's not purely self-referential; that part is fine.\n\nThe matched-capacity concern raised by the reader is real but secondary. The model adds eight oscillator families and a gate on top of the baselines; without a parameter-matched control, the gains could be extra capacity rather than chaotic response. But that issue is separate from the contradiction in the tables, which is load-bearing because it attacks the headline result directly.\n\nMy take: desk reject in current form. The internal inconsistency and duplicated rows are mechanical problems that any editor would catch, and they undermine trust in the reported results. If the authors correct the tables, resolve the ablation mismatch, and release code, it might merit a quick look in a future round. As is, there's no need to spend referee time on it.","headline":"The soft-mixture-of-oscillators idea is a reasonable incremental extension, but the paper's own tables contradict its 'consistent outperformance' claim and contain duplicated rows, so it is not referee-ready.","tokens_in":9999,"tokens_out":3342,"would_cite":false,"duration_ms":27492,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that feed-forward activation design is a crucial, underexplored factor in Transformer forecasting, and that a learnable mix of eight Lee oscillator families, gated with GELU, yields large MSE improvements on volatile…","keywords":["Time-series forecasting","Quantformer","Lee Oscillator","chaotic activation","Transformer","volatile systems","non-stationary dynamics","gated fusion"],"falsifier":"Run a matched-capacity ablation in which the eight Lee oscillator families are replaced by an equally sized mixture of trainable smooth activations (for example, eight learned GELU variants) with the same gate and training protocol; if the ETTh2 horizon-24 MSE gap over HAT disappears or reverses, the chaotic mechanism is not the causal ingredient.","tokens_in":8914,"feed_emoji":"📈","tokens_out":10298,"duration_ms":73439,"temperature":0.7,"pith_summary":"QFCQT claims that the feed-forward nonlinear transformation, not just attention, is a crucial factor in Transformer time-series forecasting. The paper proposes replacing static smooth activations with a learnable mixture of eight Lee oscillator families, compressed by Max-over-Time pooling and adaptively gated against GELU, inside a Quantformer-style linear-embedding encoder. On ETTh2 at horizon 24 it reports a 43.9% MSE improvement over HAT and 41.3% over COTN, with consistent gains across ETTh1 and an A-share stock index benchmark, supporting the claim that chaos-aware activation helps most in highly volatile regimes.","feed_headline":"Chaotic gating cuts volatile forecast error by 43.9%","feed_subtitle":"A learnable mix of eight oscillator activations replaces static Transformer feed-forward units, improving forecasts","key_machinery":"The load-bearing component is the chaotically gated feed-forward block. Each pre-activation drives $N$ internal steps of the discrete Lee oscillator dynamics $E(t+1)=f(a_1L(t)+a_2E(t)-a_3I(t)+a_4S(t)-\\xi_E)$, $I(t+1)=f(b_1L(t)-b_2E(t)-b_3I(t)+b_4S(t)-\\xi_I)$, $\\Omega(t+1)=f(S(t))$, $L(t)=[E(t)-I(t)]\\exp(-kS(t)^2)+\\Omega(t)$; the trajectory is collapsed by Max-over-Time pooling to a scalar $f^{\\mathrm{MoT}}_k(x)$. Eight such families (Table I) are mixed by softmax weights $\\pi_k$ and fused with GELU via $f_{\\mathrm{QFCQT}}(x)=\\sigma(\\lambda)\\,\\mathrm{GELU}(x)+(1-\\sigma(\\lambda)) f_{\\mathrm{chaos}}(x)$, inside a Quantformer-style linear-embedding encoder.","core_discovery":"The central discovery is that the feed-forward block's activation function can be a source of forecasting sensitivity to abrupt regime changes. QFCQT turns each scalar pre-activation into an internal oscillator trajectory across eight parameterized Lee oscillator families, pools the trajectory with Max-over-Time, mixes the families with a softmax over learnable weights, and fuses the result with GELU through a learnable scalar gate. The authors argue this preserves training stability while adding local nonlinear responsiveness, and report that the full model consistently outperforms Informer, LogTrans, LSTMa, HAT, COTN, and TimesNet across horizons, with the largest margins on the most volatile settings.","pith_inferences":["A matched-capacity control is the natural next experiment; without it, the 43.9% figure cannot be cleanly attributed to chaos rather than to the extra oscillator parameters.","The eight oscillator coefficient sets in Table I are hand-set without sensitivity analysis, so the framework's reliability across datasets and horizons may rest on those specific values.","If the mechanism is what produces the gains, then learning the oscillator parameters themselves rather than fixing them should improve or preserve performance; this is a testable extension the paper leaves open.","The 'quantum-fractal-inspired' naming is explicitly analogical, so the paper's practical contribution is the gated oscillator activation, not a new physical theory."],"forward_implications":["Transformer forecasters should be evaluated with activation design as a controlled variable, not just attention and tokenization.","Oscillator trajectories, once collapsed by Max-over-Time pooling, can serve as practical activations without destroying the parameter count of a standard feed-forward block.","A learnable smooth-chaotic gate offers a stable way to inject strong nonlinearity, letting the model decide per layer how much chaotic response to admit."],"supporting_citations":[{"why":"Supplies the Quantformer-style linear-embedding backbone that QFCQT adopts.","marker":"[2]"},{"why":"Provides the Lee oscillator activation, Max-over-Time pooling, and gated fusion pattern that QFCQT extends to eight families.","marker":"[3]"},{"why":"Defines the Lee Oscillator dynamics used in the activation module.","marker":"[4]"},{"why":"Informer is a strong Transformer baseline that QFCQT must beat in the comparisons.","marker":"[9]"},{"why":"TimesNet serves as a general time-series analysis baseline in the evaluation tables.","marker":"[13]"}],"fun_headline_variants":["QFCQT: chaotic gating of activations tames volatile series","Soft oscillator superposition beats Informer on volatile benchmarks","Eight learnable oscillator families gate each neuron's response","Activations become oscillator trajectories via Max-over-Time pooling","Chaotically gated quantformer forecast beats Transformers on volatile data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes the forecast improvements come from chaos-aware activation itself, but it never tests an equally parameterized smooth alternative, so added feed-forward capacity could explain the gains as easily as the oscillator dynamics.","fun_headline_variants_meta":{"raw":{"variants":["QFCQT: chaotic gating of activations tames volatile series","Soft oscillator superposition beats Informer on volatile benchmarks","Eight learnable oscillator families gate each neuron's response","Activations become oscillator trajectories via Max-over-Time pooling","Chaotically gated quantformer forecast beats Transformers on volatile data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001223,"raw_usage":{"total_tokens":5036,"prompt_tokens":960,"completion_tokens":4076,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":3994}},"tokens_in":576,"tokens_out":4076,"duration_ms":27278,"temperature":1.0,"reasoning_tokens":3994,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T05:30:14.517150+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a matched-capacity ablation in which the eight Lee oscillator families are replaced by an equally sized mixture of trainable smooth activations (for example, eight learned GELU variants) with the same gate and training protocol; if the ETTh2 horizon-24 MSE gap over HAT disappears or reverses, the chaotic mechanism is not the causal ingredient.","supporting_citations":[{"cited_title":"Cotn: A chaotic oscillatory transformer network for complex volatile systems under extreme conditions,","cited_arxiv_id":null,"evidence_quote":"Provides the Lee oscillator activation, Max-over-Time pooling, and gated fusion pattern that QFCQT extends to eight families."},{"cited_title":"A transient-chaotic autoassociative network (tcan) based on lee oscillators,","cited_arxiv_id":null,"evidence_quote":"Defines the Lee Oscillator dynamics used in the activation module."}],"review_version":1}