{"id":"2762b4b2-bab6-4583-852a-72767e5947c6","arxiv_id":"2608.10222","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A JEPA-style latent world model predicts the wireless propagation field in a shared representation and uses it as a prior for sparse-pilot channel reconstruction, improving symbol detection and beamforming in simulated MIMO-OFDM systems.","lead":"A machine-learning system predicts the stable parts of wireless channels in a compressed latent space instead of forecasting every complex coefficient, then combines that prediction with sparse pilot signals to reconstruct the full channel. If the results hold beyond simulation, the approach could lower channel-estimation overhead and improve beamforming in MIMO systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central cross-band beamforming claim rests on a dataset where all carrier bands share identical paths, so the JEPA spatial-subspace advantage may be an artifact of that shared-geometry assumption rather than a general property.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the synthetic cross-band data reuse a single 3.5 GHz path layout across all bands, which artificially favors latent transfer. My stress-test agrees and sharpens the point by tracing it to the precise claim in Section VIII-C, where the beamforming advantage is measured on missing carrier bands whose paths are identical to those of the observed bands. This is not a flaw in the internal logic of the JEPA pipeline; the ablation against Direct Training and the current-latent oracle is appropriate. Rather, it is an external-validity gap: the central generalization about physics-aware field representation is supported only under a shared-path geometric model that removes the main frequency-dependent physical effect that would challenge cross-band transfer. The reader already rendered a CONDITIONAL verdict, and this concern reinforces that condition rather than overturning the paper. Secondary issues, such as the absence of error bars and the unreported non-convergence of pure-prediction baselines, are real but less decisive; if the cross-band dataset is corrected, those issues should also be revisited.","tokens_in":17988,"tokens_out":5616,"duration_ms":67661,"concrete_test":"Generate a second Sionna RT dataset over the same Chicago scene with frequency-dependent material parameters and per-band path visibility, e.g., by running ray tracing independently at each of the eight carrier frequencies, or by applying a band-dependent attenuation and random path-pruning rule that removes a fraction of paths at mmWave frequencies. Keep the FWM architecture, pretraining objectives, downstream training protocol, and evaluation metrics identical, and rerun the Section VIII-C cross-band beamforming comparison for P=10 and P=4. If the FWM-versus-DT beamforming gap shrinks substantially or FWM no longer tracks the current-latent reference, then the claimed spatial-subspace preservation is specific to identical-path geometries, and the conclusion must be restricted accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that JEPA pretraining preserves the dominant spatial eigendirection and covariance structure of the missing-band channel, producing large beamforming gains even when coefficient-level NMSE is similar to an end-to-end baseline. This claim is established only in the cross-band setting of Section VIII-C, where the data-generation protocol in Section VI.A reuses the 3.5 GHz path gains, delays, and angles for all eight bands, recomputing only carrier-dependent phase, Doppler, array response, and OFDM frequency response. The paper explicitly acknowledges that it does not model frequency-dependent material responses or band-dependent path visibility. Consequently, the cross-band reconstruction problem becomes one of transferring a shared geometric support whose only band-dependent variations are phase-like, precisely the regime most favorable to a latent propagation-field prior. In real sub-6-GHz and mmWave deployments, paths appear and disappear due to frequency-selective penetration, reflection, diffraction, and blockage, and gains vary significantly across bands; under those conditions the predicted latent field may no longer align with the dominant spatial eigenspace of a missing band. Section VI.A cautions that the cross-band experiments evaluate transfer under a shared-path geometric model, but the abstract and conclusion state the stronger claim that FWM captures task-relevant spatial propagation structure and moves toward physics-aware field representation. The load-bearing gap is therefore not an internal inconsistency in the training pipeline, but a mismatch between the generality of the central conclusion and the deliberately geometry-identical simulation used to support it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes FWM, a JEPA-based latent world model for MIMO-OFDM CSI. It maps multi-resolution CSI observations from several carrier bands into a shared latent propagation-field space via scale-specific tokenizer heads, trains a latent prediction backbone with a masked/future prediction objective, and uses the predicted latent field as a structured prior for downstream channel reconstruction, either per band or across bands with missing pilots. The authors evaluate latent-prediction NMSE, multi-scale alignment strategies, and downstream reconstruction NMSE, symbol error rate, and single-stream beamforming metrics on a Sionna RT dataset generated in a Chicago city-center scenario. The central empirical claim is that JEPA pretraining preserves the dominant spatial eigendirection and covariance structure of missing-band channels, yielding substantial beamforming gains even when coefficient-level NMSE is comparable to an end-to-end baseline.","tokens_in":18220,"tokens_out":3881,"duration_ms":38074,"significance":"If the results hold beyond the specific simulation setup, the paper makes a useful contribution by shifting channel prediction from coefficient-level regression to latent predictive modeling and by demonstrating that task-relevant spatial structure can transfer across observation scales and carrier bands. The multi-scale tokenizer architecture and incremental alignment strategy are clearly formulated, and the downstream fusion of a predicted latent prior with sparse pilots is a sensible design. The beamforming evaluation metrics (eigenvector similarity, covariance NMSE, and Rayleigh quotient ratio) are appropriate and provide a credible explanation for the observed gap between reconstruction NMSE and beamforming gain. The provision of dataset and code in the footnote is a strength. However, the cross-band experiments rely on a data-generation protocol that reuses a single path layout across all bands, which severely limits the generality of the stated conclusions about physics-aware field representation.","major_comments":[{"comment":"The central cross-band beamforming claim is established only under a dataset where all eight carrier bands share exactly the same propagation paths, gains, delays, and angles, computed once at 3.5 GHz. As the authors acknowledge in Section VI.A, the dataset does not model frequency-dependent material responses or band-dependent path visibility (paths appearing or disappearing due to penetration, reflection, diffraction, or blockage). In this setting, the missing-band channel is essentially a phase-shifted, carrier-dependent version of the observed-band channel, which is precisely the regime most favorable to a latent propagation-field prior that preserves the spatial support. The conclusion in Section IX that FWM 'provides a step from channel approximation toward physics-aware field representation' therefore goes beyond what the experiment can support. To make the cross-band claim load-bearing, the authors should test the method on data where the path set itself varies across bands, for example by running Sionna RT independently at each carrier frequency or by explicitly adding/removing paths per band. Without such a test, the spatial-subspace preservation in Fig. 8 and the associated beamforming gains may be an artifact of the shared-geometry assumption.","section":"Section VI.A and Section VIII.C"},{"comment":"The statement that 'We do not report pure-prediction baselines because none converged in our experiments, including [15], [16] and the prediction-only counterpart of the proposed method' is an unsupported empirical claim that is important for positioning the paper. No convergence curves, final NMSE values, or implementation details are given for these baselines, so the reader cannot tell whether the failure is due to the dataset's phase perturbations, an implementation issue, or a fundamental property of coefficient-level prediction. Since the paper motivates JEPA precisely by arguing that raw-CSI prediction is fragile, this baseline evidence is relevant. The authors should either provide the quantitative results and learning curves for these baselines or soften the claim to an observation about their specific setup rather than a general statement about prior methods.","section":"Section VIII.A"},{"comment":"The paper asserts in the conclusion that the latent prior provides 'substantial band-wise beamforming gain even when its NMSE advantage over sufficiently trained end-to-end baselines is small or inconsistent.' However, the comparison between FWM and DT uses different training schedules (FWM downstream heads trained for 200 epochs, DT trained for 500 epochs), and the text states that the DT training schedule is longer because 'learning the complete predictor–reconstructor from scratch is more difficult.' This is a reasonable choice, but it makes the comparison less clean; the claim that JEPA is responsible for the beamforming gain would be stronger if the DTs were trained until their loss plateaued and the same number of epochs were used for the downstream stage. Please clarify whether the 500-epoch DT has converged and report its final NMSE and beamforming metrics at the same evaluation points as FWM.","section":"Section VIII.C, Fig. 9"}],"minor_comments":[{"comment":"The baseline 'SP' is mentioned ('SP and SE are retained as single-purpose baselines') but never defined; only SE is introduced in the baseline list. Please define SP or remove it.","section":"Section VIII.A"},{"comment":"The sentence 'In all cases, the allotted training schedule is sufficient for convergence' is an assertion without evidence. Presenting learning curves or final loss/epoch plots for the main JEPA and downstream training runs would make this claim verifiable.","section":"Section VI.C"},{"comment":"The downstream correspondence results in Fig. 9 are described only qualitatively in the text. Reporting numeric values for the grounding-objective ablations and for the training-strategy comparison would allow readers to judge the magnitude of the differences and would also avoid the impression that the conclusions rest on visual inspection of curves.","section":"Figure 9"},{"comment":"The statement 'The target branch is stopped from gradient back-propagation' should read 'gradient backpropagation' or 'gradient back-propagation' for consistency; also consider specifying whether the target tokenizer's batch normalization or layer normalization statistics are updated by EMA or kept frozen.","section":"Appendix B"},{"comment":"The author name 'M ´erouane' has a formatting issue with the accent; it should appear as 'Mérouane'.","section":"Author line"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of the journal and the proposed JEPA-based framework is technically sound. The main risk is that the cross-band experiment, which is the core demonstration of the paper's central claim, uses a data-generation protocol that is artificially favorable to the latent-field prior. I would encourage the editor to ask for a revision that either adds experiments with band-dependent path visibility/material responses or substantially softens the claims about physics-aware generalization. The absence of pure-prediction baseline results is also worth addressing, though it is secondary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious paper, not a hype artifact. The new idea is to predict a shared latent propagation field from multi-resolution CSI across bands, then use that predicted latent as a prior for sparse-pilot reconstruction. Prior JEPA wireless works mostly target sensing or classification; FWM is the first to use multi-resolution alignment into one field and to show downstream beamforming benefits from latent prediction. The incremental frozen-reference alignment strategy is also well designed, and the ablation showing frozen reference beats joint training is plausible.\n\nWhat earns credit: the downstream evaluation is not just NMSE. They compute SER and beamforming ratio, and they explain the mechanism—JEPA preserves the dominant spatial eigenvector and covariance structure of missing-band channels, so beamforming improves even when coefficient-level NMSE is flat. That is a real and non-obvious insight. They also include an oracle current-latent reference and ablate grounding objectives; the finding that adding raw CSI recovery hurts latent prediction is consistent with the JEPA philosophy.\n\nSoft spots: the main one is in the cross-band setup. Section VI.A states plainly that all eight bands reuse the 3.5 GHz path gains, delays, and angles, with only frequency-dependent phase, Doppler, array response, and OFDM frequency response recomputed. That makes cross-band reconstruction easier in exactly the way a latent propagation-field prior would want: the geometry is identical across bands. Real sub-6 GHz and mmWave bands do not share path visibility. The authors flag this limitation clearly, but the abstract and conclusion still lean on 'physics-aware field representation' as a general claim. That mismatch should be fixed by either adding a frequency-dependent ray-tracing dataset or softening the claim.\n\nSecond soft spot: no error bars or statistical tests anywhere. The beamforming ratio curves look convincing, but with 200 UE trajectories and one scenario, I want standard deviations or a second scene. Third: pure-prediction baselines are dismissed with 'none converged,' without curves or details. That is testable and should be shown; otherwise it reads as a convenient exclusion.\n\nThese are fixable, not fatal. The central logic—predict stable structure in latent space, fuse with sparse pilots—holds up on this evidence. The code and dataset are promised in the paper; if they ship, that strengthens it.\n\nBottom line: worth serious peer review. A good referee should push on cross-band realism and error bars; the paper will survive and be more useful. I would cite it and bring it to a reading group.","headline":"A genuinely useful JEPA-for-wireless paper with a clean latent-field idea and an honest experimental write-up; the cross-band results are real but rest on a shared-geometry simulation that should be disclosed more prominently.","tokens_in":18849,"tokens_out":1986,"would_cite":true,"duration_ms":21517,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"JEPA-based latent-field prediction preserves the spatial structure beamforming needs.","keywords":["channel state information","channel prediction","channel reconstruction","JEPA","world model","MIMO-OFDM","beamforming","cross-band transfer"],"falsifier":"Regenerate or measure the scenario with band-dependent path visibility — ray-tracing each carrier band independently so paths can appear or disappear with frequency, or using measured sub-6-GHz and mmWave channel pairs — and compare FWM against the end-to-end baseline on bands without pilots; if the beamforming gain and dominant-eigenvector similarity fall to baseline levels, the claim that the latent field transfers shared spatial structure across bands is falsified.","tokens_in":17748,"feed_emoji":"📡","tokens_out":8060,"duration_ms":69571,"temperature":0.7,"pith_summary":"This paper tries to establish that wireless channel prediction is more reliable and more useful when done in a learned latent propagation-field space rather than at the level of raw complex coefficients. The authors build a JEPA-based world model that maps multi-resolution CSI observations across several carrier bands into a shared latent field, predicts masked or future field states there, and then fuses the predicted latent with sparse current pilots to reconstruct the channel. The claimed payoff is concrete: improved symbol detection and beamforming performance close to an oracle that uses the true current latent, even when coefficient-level NMSE improvements over an end-to-end baseline are small. A reader should care because the approach separates predictable propagation structure from phase-sensitive detail, which points toward lower pilot overhead and transfer of channel knowledge across frequency bands.","feed_headline":"Latent-space channel prediction preserves beamforming gains","feed_subtitle":"A JEPA world model fuses predicted propagation structure with sparse pilots, preserving the spatial subspace beamforming needs.","key_machinery":"The central object is the latent propagation-field representation: a shared latent space in which each time–carrier-band cell is encoded as a small set of tokens by scale-specific tokenizer heads, with a shared JEPA backbone predicting masked or future cells. The 'field layer' means propagation-level structure — delay-angle support, dominant paths, spatial covariance, and dominant eigenspaces — rather than individual complex coefficients. The training objective combines latent prediction with a diversity regularizer and uses an exponential-moving-average target tokenizer, so the model encodes what is predictable across time and band while avoiding collapsed representations. The load-bearing identity is that beamforming quality follows the Rayleigh quotient of the true channel covariance evaluated at the reconstructed beamformer, so preserving the dominant eigenvector and covariance of the latent field converts directly into beamforming gain. A frozen-reference incremental alignment procedure lets new observation scales be added without retraining the shared backbone.","core_discovery":"The central claim is that a shared latent propagation state, learned by JEPA-style masked prediction, carries the spatial structure that matters for communication tasks, and that this structure survives prediction even when individual CSI coefficients do not. On a ray-traced city scenario with five sub-6-GHz and three mmWave carrier bands, the paper's FWM gives each observation scale its own tokenizer head, maps all scales into a common latent field, predicts the current-time field from history, and uses the predicted latent as a structured prior for a reconstruction head that also takes sparse noisy pilots. In the cross-band setting, where some carrier bands have no pilots, the predicted latent is calibrated using pilot-conditioned latents from the observed bands. The paper reports that FWM and a fully end-to-end trained baseline show comparable reconstruction NMSE, yet FWM retains higher dominant-eigenvector similarity and lower covariance NMSE, yielding band-wise beamforming ratios close to a current-latent oracle. The authors conclude that JEPA pretraining better preserves the dominant spatial eigendirection and covariance structure of the missing-band channel, which govern the beamforming Rayleigh quotient.","pith_inferences":["Our inference: if propagation-geometry preservation is the operative mechanism, then the FWM latent should also support other geometry-driven tasks such as localization, sensing, and beam management; the paper lists these as future work rather than demonstrated results.","Our inference: the cross-band gains are likely to shrink under frequency-dependent path visibility, because the dataset deliberately reuses one path set across all bands; a direct test is to regenerate channels with per-band ray tracing or measured dual-band data and compare beamforming ratios.","Our inference: the evaluation principle here — judge predicted CSI by what it does to beamforming and detection, not by NMSE — could reasonably be adopted by the wider channel-prediction literature, but the paper only demonstrates it in its own scenario."],"forward_implications":["Fewer pilots are needed for the same reconstruction quality, because the predicted latent field already supplies the spatial subspace that sparse pilots would otherwise have to recover.","Carrier bands with no current pilots can still be reconstructed from observed bands plus the latent prior, which bears directly on FDD-style reciprocity problems.","Coefficient-level NMSE is a poor proxy for task value: models with similar NMSE can differ sharply in beamforming gain, so communication-level metrics should accompany CSI metrics.","New observation formats can be introduced incrementally with a frozen reference branch, avoiding retraining the whole model when resolution or bandwidth changes.","Adding raw-CSI grounding objectives to the latent target is counterproductive for prediction; diversity-based JEPA pretraining outperforms direct CSI, amplitude, and diffusion grounding."],"supporting_citations":[{"why":"Introduces joint-embedding predictive architecture: predict target embeddings in latent space instead of reconstructing raw inputs.","marker":"[3]"},{"why":"Extends JEPA to video and world-model learning, motivating temporal latent prediction.","marker":"[4]"},{"why":"Shows neural networks can learn compact CSI representations for feedback, grounding the latent-representation approach.","marker":"[8]"},{"why":"Pure ODE-based channel-prediction baseline that the paper says did not converge under its perturbed scenario.","marker":"[15]"},{"why":"Robustness-focused CSI prediction baseline that also did not converge, motivating fusion of prediction with current estimation.","marker":"[16]"},{"why":"FDD cross-band reciprocity method whose problem formulation is structurally similar to the cross-band alignment setting.","marker":"[18]"},{"why":"Closest unified space-time-frequency channel prediction and reconstruction foundation model, used as comparison.","marker":"[23]"}],"fun_headline_variants":["Latent-space channel prediction retains beamforming gains","JEPA world model preserves spatial structure for beamforming","Shared latent state beats coefficient fitting for beamforming","Predict propagation in latent space to keep beamforming subspace","Cross-band channel prediction via JEPA latent fields"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that all eight simulated carrier bands see exactly the same propagation paths, gains, delays, and angles, so the latent field learned from one band is genuinely shared by the others; if real frequency-dependent effects make paths appear or disappear across bands, the cross-band transfer gains may not survive.","fun_headline_variants_meta":{"raw":{"variants":["Latent-space channel prediction retains beamforming gains","JEPA world model preserves spatial structure for beamforming","Shared latent state beats coefficient fitting for beamforming","Predict propagation in latent space to keep beamforming subspace","Cross-band channel prediction via JEPA latent fields"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000244,"raw_usage":{"total_tokens":1563,"prompt_tokens":1007,"completion_tokens":556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":623,"completion_tokens_details":{"reasoning_tokens":483}},"tokens_in":623,"tokens_out":556,"duration_ms":5290,"temperature":1.0,"reasoning_tokens":483,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:11:00.978581+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regenerate or measure the scenario with band-dependent path visibility — ray-tracing each carrier band independently so paths can appear or disappear with frequency, or using measured sub-6-GHz and mmWave channel pairs — and compare FWM against the end-to-end baseline on bands without pilots; if the beamforming gain and dominant-eigenvector similarity fall to baseline levels, the claim that the latent field transfers shared spatial structure across bands is falsified.","supporting_citations":[{"cited_title":"ODE-Former for mobile channel prediction: A novel learning structure leveraging the physics continuity,","cited_arxiv_id":null,"evidence_quote":"Pure ODE-based channel-prediction baseline that the paper says did not converge under its perturbed scenario."},{"cited_title":"CSI-4CAST: A hybrid deep learning model for CSI prediction with comprehensive robustness and generalization testing,","cited_arxiv_id":null,"evidence_quote":"Robustness-focused CSI prediction baseline that also did not converge, motivating fusion of prediction with current estimation."},{"cited_title":"FIRE: Enabling reciprocity for FDD MIMO systems,","cited_arxiv_id":null,"evidence_quote":"FDD cross-band reciprocity method whose problem formulation is structurally similar to the cross-band alignment setting."},{"cited_title":"WiFo: Wireless foundation model for channel prediction,","cited_arxiv_id":null,"evidence_quote":"Closest unified space-time-frequency channel prediction and reconstruction foundation model, used as comparison."}],"review_version":1}