REVIEW 2 major objections 5 minor 33 references
CFM-Bench provides a common substrate for comparing channel foundation models across six radio configurations and six task groups, using leakage-resistant partitions and a test-exposure policy that make transfer comparisons trustworthy.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 00:31 UTC pith:UINJPSK5
load-bearing objection A coherent, genuinely useful benchmark protocol for CFMs, but it ships without baseline experiments or usable links, and its leakage guarantee is honor-system only. the 2 major comments →
CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that CFM-Bench makes cross-model comparison meaningful by fixing the things that currently vary between papers. It selects one fixed radio configuration per source, partitions at the largest independent physical unit (complete trajectories, measurement sessions, vehicle links, simulation realizations, or buffered spatial regions), and forbids any benchmark split from being used in foundation-model pretraining. It also requires a data-exposure statement listing every dataset used during development, and disables scientifically unsupported task–domain combinations rather than manufacturing labels. The result is a shared substrate on which a pretrained channel representatio
What carries the argument
The load-bearing mechanism is unit-level leakage-resistant partitioning combined with a mandatory data-exposure policy and a task-support matrix. Partitions are drawn at the level of complete physical units so that spatially or temporally correlated samples never straddle the train/test boundary; the policy reserves official splits exclusively for fine-tuning and scoring; and the task-support matrix encodes which tasks are physically meaningful per domain, preventing superficially similar labels from being compared under incompatible semantics.
Load-bearing premise
The fairness guarantee rests on voluntary disclosure and public test sets; if a participant silently uses test units during development, the benchmark's central promise of trustworthy comparison collapses.
What would settle it
Compute the average complex-CSI similarity between official training and test units and compare it with the similarity within the training set. If the cross-split similarity distribution substantially overlaps the within-split distribution, the unit-level partitioning has not removed information leakage, and rankings built on the benchmark would be inflated.
If this is right
- Any pretrained channel model can be ranked against other CFMs and against task-specific networks under identical data, splits, and metrics.
- Reported pretraining gains can be checked for authenticity: gains that vanish under unit-level isolation are exposed as leakage artifacts.
- Transferability can be assessed across statistical, ray-traced, measured, and multimodal channels within a single protocol.
- Per-domain scores become the unit of comparison, with an unweighted macro-average explicitly demoted to a secondary summary.
- Researchers get a fixed test-exposure policy that disambiguates compliant results from test-exposed or transductive ones.
Where Pith is reading between the lines
- The benchmark's design suggests a natural next step: adding a hidden test-set tier would close the acknowledged gap that public test sets cannot prevent repeated manual adaptation.
- The task-support matrix—disabling unsupported domain–task combinations—could become a template for other foundation-model benchmarks where physical semantics vary by domain.
- Because domains differ in difficulty and sample count, the macro-average score should be read with caution; per-domain inspection will likely be more informative than any single number.
- The strict exclusion of tasks like temporal extrapolation on measured domains may understate what sophisticated signal processing can extract; future releases could add processed variants as separate tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. CFM-Bench is a benchmark/resource paper for channel foundation models (CFMs). It curates one fixed radio configuration from each of six public channel data sources — a 3GPP statistical urban-microcell domain, two ray-tracing domains (Wireless InSite and Sionna/MOCSID), two measured massive-MIMO domains (DICHASUS, MaMIMO-UAV), and a synchronized vehicular multimodal domain — and imposes unit-level train/validation/test partitions based on trajectories, sessions, vehicle links, simulators, or spatial regions. The paper defines six task groups spanning PHY, RAN, and ISAC, with per-domain eligibility rules, metrics such as NMSE, SGCS, Macro-F1, Top-k beam accuracy, and localization error, and a mandatory test-exposure/data-disclosure policy. The central claim is that the benchmark provides a common substrate for fair and trustworthy comparison of CFMs across models, domains, and tasks.
Significance. If adopted, CFM-Bench would address a real gap in CFM evaluation: the lack of a unified protocol with matched downstream tasks and leakage-resistant partitions. The paper has several genuine strengths: it spans complementary channel-generation mechanisms, disables unsupported domain-task combinations instead of forcing labels, retains physical metadata without prescribing a fixed input shape, provides explicit per-domain metrics and codebook definitions, and documents quality-control and licensing choices. I found no circular derivation or hidden fitted parameters; this is a resource paper. However, the paper's central promise that its test-isolation policy prevents undeclared test-set reuse is not currently enforceable with a fully public test set and self-reported disclosure, and no baseline experiments demonstrate that the proposed tasks and partitions behave as intended. Both issues are fixable, but they are load-bearing for the benchmark's fairness and usability claims.
major comments (2)
- [Sec. V.A and Sec. VII] The test-isolation guarantee is not operational. All test units are released publicly, there is no hidden evaluation server, and enforcement rests solely on a mandatory data-exposure statement. Because the six domains derive from public upstream datasets (DeepMIMO, MOCSID, DICHASUS, MaMIMO-UAV, Multimodal-Wireless), a participant can obtain the same held-out trajectories, sessions, flights, or vehicle links from the original repositories without touching CFM-Bench files, making any detection impossible. Section VII itself concedes: 'The public test sets favor reproducibility but cannot prevent undeclared reuse or repeated manual adaptation.' This concession contradicts the Abstract's promise to 'prevent pretraining leakage' and contribution bullet 3's claim that the test-isolation rule 'ensures' a test-exposed model cannot be presented as compliant. The fairness claim is therefore condit
- [Secs. IV-V, Tables II and V] The benchmark defines official splits, tasks, and metrics but reports no experimental validation. There are no baselines showing that any of the six task groups is solvable, that the official metrics produce meaningful and stable values, or that unit-level partitions create a measurable train/test gap. For example, Section VII states that E2 future-beam prediction 'admits a strong persistence baseline,' yet no persistence baseline is reported; M1 localization permits RGB/LiDAR inputs that can reveal absolute position through visual landmarks, but no modality ablation is provided to show whether the task measures channel representations or visual place recognition. Without at least simple baselines (random/prior, linear models, small neural networks, persistence for temporal tasks) and a demonstration that performance degrades on held-out units relative to random splits, the claims of 'st
minor comments (5)
- [Eq. (2)] N in the SGCS formula is not defined in the text. It presumably denotes the number of samples; please state this explicitly.
- [Table I, R2 row] The row lists 'Unspecified / 1.92 MHz' for carrier/bandwidth, while Sec. IV.B defines a derived 64-tone, 30-kHz relative baseband grid. Please clarify the relation between the upstream dataset's bandwidth and the benchmark-defined grid, and state whether the 1.92 MHz figure is from the original MOCSID release.
- [Sec. V.E] M1 localization prohibits pose, GPS, and world-coordinate fields, but allows RGB and LiDAR. Since these modalities can reveal absolute position through visual landmarks, please state whether a CSI-only ranking will be maintained or explicitly report modality-controlled baselines. Otherwise the channel-model interpretation of the M1 score is ambiguous.
- [Sec. IV.E] The temporal test views are described by number of windows and window lengths, but it is not specified whether scores are computed per window, per frame, or aggregated across windows. Please define the official aggregation for temporal tasks.
- [Sec. VI] The paper states that evaluation software, split definitions, and documentation are released, but it provides no repository URL, DOI, or persistent identifier for the benchmark release itself. Please add one.
Circularity Check
No circularity: CFM-Bench is a benchmark/resource paper with no fitted parameter renamed as a prediction, no uniqueness claim imported from authors, and no load-bearing self-citation chain.
full rationale
CFM-Bench does not present a derivation chain that reduces to its inputs. It curates six existing data sources, defines partitions, task protocols, and metrics, and releases them as a benchmark substrate. There is no fitted parameter that is later called a prediction; the benchmark's claims are about providing evaluation infrastructure, not about deriving empirical results from a theory. Self-citations appear only as contextual references (e.g., [4] for the CFM concept, [18] for CSI-CLIP++, [27]-[33] for surveys) and are not used to justify the benchmark's validity or to force a modeling choice. The paper explicitly concedes its central enforcement limitation in Section VII: 'The public test sets favor reproducibility but cannot prevent undeclared reuse or repeated manual adaptation.' That is an acknowledged limitation of the benchmark's fairness guarantee, not a circular step: the argument does not assume the conclusion, and no equation or fitted value is equivalent to the input by construction. Consequently, no specific circular step can be quoted, and the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Unit-level isolation is sufficient to prevent information leakage between train and test.
- domain assumption Self-disclosed data-exposure statements ensure test isolation in practice.
- domain assumption The six selected configurations are representative of channel diversity for benchmarking CFMs.
read the original abstract
Channel foundation models (CFMs) are developing rapidly, with recent studies reporting benefits from pretraining across downstream wireless tasks. Yet CFMs are commonly evaluated in model-specific pipelines with different data, radio configurations, partitions, adaptation procedures, task definitions, and metrics. Reported comparisons therefore tend to show that pretraining improves over supervised training from scratch within one pipeline, but neither rank CFMs nor compare them fairly with task-specific models. We release CFM-Bench, a unified multi-domain, multi-task benchmark designed to address this gap. It curates six channel configurations spanning 3GPP statistical simulation, two independent ray-tracing pipelines, industrial and aerial measurements, and synchronized vehicular multimodal simulation. Official partitions isolate complete trajectories, measurement sessions, vehicle links, simulation realizations, or buffered spatial regions. CFM-Bench does not prescribe an external pretraining corpus or strategy; no benchmark split may be used for foundation-model pretraining, and the official training split is reserved exclusively for downstream fine-tuning. The benchmark additionally requires disclosure of all data used during model development and prohibits training-stage use of official test units. Six task groups are organized along three CFM application dimensions: physical-layer (PHY) channel intelligence, radio-access-network (RAN) decision intelligence, and integrated sensing and communication (ISAC). They cover CSI feedback, frequency and temporal channel extrapolation, propagation-state classification, current- and future-beam prediction, and single-frame and temporal localization. CFM-Bench provides a common substrate for comparing the transferability of channel representations across models, domains, and tasks.
Figures
Reference graph
Works this paper leans on
-
[1]
Study on artificial intelligence (AI)/machine learning (ML) for NR air interface,
3GPP, “Study on artificial intelligence (AI)/machine learning (ML) for NR air interface,” 3rd Generation Partnership Project (3GPP), Technical Report TR 38.843 V18.0.0, Dec. 2023
2023
-
[2]
Ai/ml for mobile networks: Current status in rel. 19 and challenges ahead,
Y . Gao, X. Wu, J. Jiang, B. Hu, J. Du, Q. Ye, S. Zhang, F. R. Yu, and S. Xu, “Ai/ml for mobile networks: Current status in rel. 19 and challenges ahead,”arXiv preprint arXiv:2603.14317, 2026
arXiv 2026
-
[3]
On the opportunities and risks of foundation models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskillet al., “On the opportunities and risks of foundation models,”arXiv preprint arXiv:2108.07258, 2021
Pith/arXiv arXiv 2021
-
[4]
Towards channel foundation models (CFMs): Motivations, methodologies and opportunities,
J. Jiang, Y . Gao, X. Wu, and S. Xu, “Towards channel foundation models (CFMs): Motivations, methodologies and opportunities,”arXiv preprint arXiv:2507.13637, 2025
Pith/arXiv arXiv 2025
-
[5]
Lwm: A pre-trained wire- less foundation model for universal feature extraction,
S. Alikhani, G. Charan, and A. Alkhateeb, “Lwm: A pre-trained wire- less foundation model for universal feature extraction,” in2025 IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN), 2025, pp. 1–6
2025
-
[6]
WiFo: Wireless foundation model for channel prediction,
B. Liu, S. Gao, X. Liu, X. Cheng, and L. Yang, “WiFo: Wireless foundation model for channel prediction,”Sci. China Inf. Sci., vol. 68, no. 6, p. 162302, 2025
2025
-
[7]
Csi-mae: A masked autoencoder- based channel foundation model,
J. Jiang, X. Ruan, and S. Xu, “Csi-mae: A masked autoencoder- based channel foundation model,” 2026. [Online]. Available: https: //arxiv.org/abs/2601.03789
arXiv 2026
-
[8]
DeepMIMO: A generic deep learning dataset for mil- limeter wave and massive MIMO applications,
A. Alkhateeb, “DeepMIMO: A generic deep learning dataset for mil- limeter wave and massive MIMO applications,” in2019 Information Theory and Applications Workshop (ITA), San Diego, CA, USA, Feb. 2019, pp. 1–8
2019
-
[9]
Sionna: An open-source library for next-generation physical layer research,
J. Hoydis, S. Cammerer, F. A ¨ıt Aoudia, A. Vem, N. Binder, G. Marcus, and A. Keller, “Sionna: An open-source library for next-generation physical layer research,”arXiv preprint arXiv:2203.11854, 2022
Pith/arXiv arXiv 2022
-
[10]
Sionna RT: Differentiable ray tracing for radio propagation modeling,
J. Hoydis, F. A ¨ıt Aoudia, S. Cammerer, M. Nimier-David, N. Binder, G. Marcus, and A. Keller, “Sionna RT: Differentiable ray tracing for radio propagation modeling,” in2023 IEEE Globecom Workshops (GC Wkshps), Dec. 2023, pp. 317–321
2023
-
[11]
Multi-cell outdoor channel state information dataset (MOCSID),
M. E. M. Makhlouf, M. Guillaud, and Y . Vindas, “Multi-cell outdoor channel state information dataset (MOCSID),” in2025 Joint European Conference on Networks and Communications and 6G Summit (Eu- CNC/6G Summit), Jun. 2025, pp. 85–90
2025
-
[12]
Multi-cell outdoors channel state information dataset (MOCSID),
——, “Multi-cell outdoors channel state information dataset (MOCSID),” Feb. 2025. [Online]. Available: https://doi.org/10.5281/ zenodo.14535165
2025
-
[13]
A distributed massive MIMO channel sounder for “big CSI data
F. Euchner, M. Gauger, S. D ¨orner, and S. ten Brink, “A distributed massive MIMO channel sounder for “big CSI data”-driven machine learning,” inWSA 2021; 25th International ITG Workshop on Smart Antennas, 2021, pp. 1–6
2021
-
[14]
CSI dataset dichasus-adxx: ARENA2036: Distributed setup in industrial environment at 3.4 GHz,
F. Euchner, P. Stephan, M. Gauger, and S. ten Brink, “CSI dataset dichasus-adxx: ARENA2036: Distributed setup in industrial environment at 3.4 GHz,” 2024. [Online]. Available: https://doi.org/10. 18419/darus-4062
2024
-
[15]
CSI measurements and initial results for massive MIMO to UA V communication,
Z. Cui, A. Colpaert, and S. Pollin, “CSI measurements and initial results for massive MIMO to UA V communication,” in2023 57th Asilomar Conference on Signals, Systems, and Computers, Oct. 2023
2023
-
[16]
MaMIMO-UA V 3d channel state information dataset,
A. Colpaert, C. Thys, Z. Cui, and S. Pollin, “MaMIMO-UA V 3d channel state information dataset,” 2023. [Online]. Available: https://doi.org/10.48804/0IMQDF
-
[17]
Multimodal- wireless: A large-scale dataset for sensing and communication,
T. Mao, L. Liang, J. Yang, H. Ye, S. Jin, and G. Y . Li, “Multimodal- wireless: A large-scale dataset for sensing and communication,” in 2026 IEEE International Conference on Communications (ICC), 2026. [Online]. Available: https://arxiv.org/abs/2511.03220
arXiv 2026
-
[18]
CSI-CLIP++: A scalable channel foundation model for wireless communication via CIR–CSI consistency,
J. Jiang, W. Yu, Y . Li, Y . Gao, and S. Xu, “CSI-CLIP++: A scalable channel foundation model for wireless communication via CIR–CSI consistency,”arXiv preprint arXiv:2606.25714, 2026
Pith/arXiv arXiv 2026
-
[19]
Y . Wang, L. Sun, T. Yang, Y . Shi, M. Elkashlan, and X. Tang, “Filter-and-attend: Wireless channel foundation model with noise-plus- interference suppression structure,”arXiv preprint arXiv:2509.15993, 2026
arXiv 2026
-
[20]
A wireless foundation model for multi-task prediction,
Y . Sheng, J. Wang, X. Zhou, L. Liang, H. Ye, S. Jin, and G. Y . Li, “A wireless foundation model for multi-task prediction,”arXiv preprint arXiv:2507.05938, 2025
Pith/arXiv arXiv 2025
-
[21]
WiFo-E: A scalable wireless foundation model for end-to-end FDD precoding in communication networks,
W. Wen, S. Gao, H. Zhang, X. Cheng, and L. Yang, “WiFo-E: A scalable wireless foundation model for end-to-end FDD precoding in communication networks,”arXiv preprint arXiv:2601.09186, 2026
arXiv 2026
-
[22]
Study on channel model for frequencies from 0.5 to 100 GHz,
3GPP, “Study on channel model for frequencies from 0.5 to 100 GHz,” 3rd Generation Partnership Project (3GPP), Technical Report TR 38.901 V18.0.0, May 2024
2024
-
[23]
Overview of deep learning- based CSI feedback in massive MIMO systems,
J. Guo, C.-K. Wen, S. Jin, and G. Y . Li, “Overview of deep learning- based CSI feedback in massive MIMO systems,”IEEE Transactions on Communications, vol. 70, no. 12, pp. 8017–8045, Dec. 2022
2022
-
[24]
Overview of deep learning-based CSI feedback in massive MIMO systems,
——, “Overview of deep learning-based CSI feedback in massive MIMO systems,”IEEE Transactions on Communications, vol. 70, no. 12, pp. 8017–8045, 2022. 10
2022
-
[25]
Study on artificial intelligence (AI)/machine learning (ML) for NR air interface (Release 18),
3GPP, “Study on artificial intelligence (AI)/machine learning (ML) for NR air interface (Release 18),” 3rd Generation Partnership Project (3GPP), Tech. Rep. TR 38.843 V18.0.0, Dec. 2023
2023
-
[26]
AI for CSI feedback enhancement in 5G-Advanced,
J. Guo, C.-K. Wen, S. Jin, and X. Li, “AI for CSI feedback enhancement in 5G-Advanced,”IEEE Wireless Communications, vol. 31, no. 3, pp. 169–176, 2024
2024
-
[27]
AI-driven channel state information (CSI) extrapolation for 6G: Current situations, challenges and future research,
Y . Gao, Z. Lu, X. Wu, W. Yu, S. Liu, J. Du, S. Zhang, X. Chu, and S. Xu, “AI-driven channel state information (CSI) extrapolation for 6G: Current situations, challenges and future research,”IEEE Communica- tions Surveys & Tutorials, vol. 28, pp. 4485–4518, Jan. 2026
2026
-
[28]
SSNet: Flexible and robust channel extrapolation for fluid antenna sys- tems enabled by an self-supervised learning framework,
Y . Gao, Y . Liu, R. Yu, S. Liu, Y . Jin, S. Zhang, S. Xu, and X. Chu, “SSNet: Flexible and robust channel extrapolation for fluid antenna sys- tems enabled by an self-supervised learning framework,”IEEE Journal on Selected Areas in Communications, vol. 44, pp. 1276–1289, 2026
2026
-
[29]
Enabling 6g through multi-domain channel extrapolation: Opportunities and challenges of generative artificial intelligence,
Y . Gao, Z. Lu, Y . Wu, Y . Jin, S. Zhang, X. Chu, S. Xu, and C.- X. Wang, “Enabling 6g through multi-domain channel extrapolation: Opportunities and challenges of generative artificial intelligence,”IEEE Communications Magazine, vol. 64, no. 1, pp. 222–228, 2026
2026
-
[30]
Generalizable and robust beam prediction for 6g networks: An deep- learning framework with positioning feature fusion,
Y . Jin, Y . Li, J. Jun, Y . Gao, S. Liu, J. Du, Z. Yang, and S. Xu, “Generalizable and robust beam prediction for 6g networks: An deep- learning framework with positioning feature fusion,”IEEE Transactions on Network Science and Engineering, 2026
2026
-
[31]
A survey of beam management for mmWave and THz communications towards 6G,
Q. Xue, C. Ji, S. Ma, J. Guo, Y . Xu, Q. Chen, and W. Zhang, “A survey of beam management for mmWave and THz communications towards 6G,”IEEE Communications Surveys & Tutorials, vol. 26, no. 3, pp. 1520–1559, 2024
2024
-
[32]
Sidelink positioning: Standardization advancements, challenges and opportunities,
Y . Gao, G. Pan, Z. Zhong, Z. Jinm, Y . Hu, Y . Jin, , and S. Xu, “Sidelink positioning: Standardization advancements, challenges and opportunities,”IEEE Communications Magazine, vol. 64, no. 4, pp. 128– 134, 2026
2026
-
[33]
Enhanced fingerprint-based positioning with practical imperfections: Deep learning-based approaches,
S. Xu, J. Jiang, W. Yu, Y . Gao, G. Pan, S. Mu, Z. Ai, Y . Gao, P. Jiang, and C.-X. Wang, “Enhanced fingerprint-based positioning with practical imperfections: Deep learning-based approaches,”IEEE Wireless Communications, vol. 33, no. 1, pp. 252–258, Jan. 2026
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.