{"id":"8b2cbbcf-944b-44f2-a7c7-7593981e0678","arxiv_id":"2607.14975","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CFM-Bench provides six curated radio-channel datasets with leakage-resistant splits and six task groups for standardizing evaluation of channel foundation models.","lead":"CFM-Bench is a new benchmark for comparing 'channel foundation models' — AI models trained on radio signals — across six different kinds of wireless data. It fixes data splits, task definitions, and rules against data leakage so that different models can be ranked fairly against each other and against task-specific baselines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Public test sets plus voluntary disclosure cannot enforce the test-isolation rule; an undeclared participant can pretrain on test units and distort rankings, so the benchmark's fair-comparison guarantee is not currently operational.","rationale":"The reader's weakest assumption is precisely that the test-isolation guarantee rests on voluntary disclosure and public test sets; the manuscript's own Section VII says this cannot prevent undeclared reuse. I agree this is the load-bearing point. The benchmark is thoughtfully designed otherwise: unit-level splits are defined per source, task support is grounded in data semantics, and limitations are candidly stated. But its central promise—'preventing pretraining leakage' and enabling a trustworthy ranking—requires enforcement, not just a rule. No current benchmark can stop a determined cheater with public test sets; however, the paper could mitigate the concern by adding a hidden audit subset or an API-gated test set, and by including the code/data links and a small baseline suite. Until then, the fair-comparison claim is conditional. Because the reader already rendered CONDITIONAL, my assessment does not move the verdict; it reinforces it. I set UNCHANGED.","tokens_in":13103,"tokens_out":6584,"duration_ms":77399,"concrete_test":"Run a controlled contamination probe on one domain (e.g., R2 or E2): train the same downstream model twice—once on the official training split only, once additionally on the official test units (simulating an undeclared cheater). Submit both under the current protocol and check whether any automated part of the release/evaluation pipeline flags the contaminated model. If the contaminated model's scores improve and no signal is raised, the test-isolation guarantee is not operational; adding a hidden audit subset (e.g., 20% of test units withheld from public release and scored only by the organizers) would settle whether the public-test ranking remains trustworthy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"CFM-Bench's central value is a trustworthy common substrate for ranking CFMs. This depends on the claim (Sec. I, contribution bullet 3) that a strict test-isolation rule means 'any model that touched an official test unit during development cannot be presented as a compliant result.' But the enforcement mechanism is only a mandatory data-exposure statement, and all test units are publicly released; there is no hidden server or cryptographic holdout. Section VII concedes exactly this: 'The public test sets favor reproducibility but cannot prevent undeclared reuse or repeated manual adaptation.' A participant who silently pretrains on test trajectories/sessions/links will obtain scores that are not comparable to honest submissions, and the benchmark has no way to detect or disqualify them. Because the domains are built from public upstream datasets (DeepMIMO, MOCSID, DICHASUS, MaMIMO-UAV, Multimodal-Wireless), a participant can also obtain adjacent scene data from the original source that overlaps test units without touching the benchmark's exact files, defeating even unit-level isolation. The paper's proposed future hidden leaderboard is not part of the current release. Thus the central fair-comparison claim is conditional on participant honesty, which is exactly the condition a benchmark is supposed to remove.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"CFM-Bench is a benchmark/resource paper for channel foundation models (CFMs). It curates one fixed radio configuration from each of six public channel data sources — a 3GPP statistical urban-microcell domain, two ray-tracing domains (Wireless InSite and Sionna/MOCSID), two measured massive-MIMO domains (DICHASUS, MaMIMO-UAV), and a synchronized vehicular multimodal domain — and imposes unit-level train/validation/test partitions based on trajectories, sessions, vehicle links, simulators, or spatial regions. The paper defines six task groups spanning PHY, RAN, and ISAC, with per-domain eligibility rules, metrics such as NMSE, SGCS, Macro-F1, Top-k beam accuracy, and localization error, and a mandatory test-exposure/data-disclosure policy. The central claim is that the benchmark provides a common substrate for fair and trustworthy comparison of CFMs across models, domains, and tasks.","tokens_in":13405,"tokens_out":8266,"duration_ms":102239,"significance":"If adopted, CFM-Bench would address a real gap in CFM evaluation: the lack of a unified protocol with matched downstream tasks and leakage-resistant partitions. The paper has several genuine strengths: it spans complementary channel-generation mechanisms, disables unsupported domain-task combinations instead of forcing labels, retains physical metadata without prescribing a fixed input shape, provides explicit per-domain metrics and codebook definitions, and documents quality-control and licensing choices. I found no circular derivation or hidden fitted parameters; this is a resource paper. However, the paper's central promise that its test-isolation policy prevents undeclared test-set reuse is not currently enforceable with a fully public test set and self-reported disclosure, and no baseline experiments demonstrate that the proposed tasks and partitions behave as intended. Both issues are fixable, but they are load-bearing for the benchmark's fairness and usability claims.","major_comments":[{"comment":"The test-isolation guarantee is not operational. All test units are released publicly, there is no hidden evaluation server, and enforcement rests solely on a mandatory data-exposure statement. Because the six domains derive from public upstream datasets (DeepMIMO, MOCSID, DICHASUS, MaMIMO-UAV, Multimodal-Wireless), a participant can obtain the same held-out trajectories, sessions, flights, or vehicle links from the original repositories without touching CFM-Bench files, making any detection impossible. Section VII itself concedes: 'The public test sets favor reproducibility but cannot prevent undeclared reuse or repeated manual adaptation.' This concession contradicts the Abstract's promise to 'prevent pretraining leakage' and contribution bullet 3's claim that the test-isolation rule 'ensures' a test-exposed model cannot be presented as compliant. The fairness claim is therefore condit","section":"Sec. V.A and Sec. VII"},{"comment":"The benchmark defines official splits, tasks, and metrics but reports no experimental validation. There are no baselines showing that any of the six task groups is solvable, that the official metrics produce meaningful and stable values, or that unit-level partitions create a measurable train/test gap. For example, Section VII states that E2 future-beam prediction 'admits a strong persistence baseline,' yet no persistence baseline is reported; M1 localization permits RGB/LiDAR inputs that can reveal absolute position through visual landmarks, but no modality ablation is provided to show whether the task measures channel representations or visual place recognition. Without at least simple baselines (random/prior, linear models, small neural networks, persistence for temporal tasks) and a demonstration that performance degrades on held-out units relative to random splits, the claims of 'st","section":"Secs. IV-V, Tables II and V"}],"minor_comments":[{"comment":"N in the SGCS formula is not defined in the text. It presumably denotes the number of samples; please state this explicitly.","section":"Eq. (2)"},{"comment":"The row lists 'Unspecified / 1.92 MHz' for carrier/bandwidth, while Sec. IV.B defines a derived 64-tone, 30-kHz relative baseband grid. Please clarify the relation between the upstream dataset's bandwidth and the benchmark-defined grid, and state whether the 1.92 MHz figure is from the original MOCSID release.","section":"Table I, R2 row"},{"comment":"M1 localization prohibits pose, GPS, and world-coordinate fields, but allows RGB and LiDAR. Since these modalities can reveal absolute position through visual landmarks, please state whether a CSI-only ranking will be maintained or explicitly report modality-controlled baselines. Otherwise the channel-model interpretation of the M1 score is ambiguous.","section":"Sec. V.E"},{"comment":"The temporal test views are described by number of windows and window lengths, but it is not specified whether scores are computed per window, per frame, or aggregated across windows. Please define the official aggregation for temporal tasks.","section":"Sec. IV.E"},{"comment":"The paper states that evaluation software, split definitions, and documentation are released, but it provides no repository URL, DOI, or persistent identifier for the benchmark release itself. Please add one.","section":"Sec. VI"}],"recommendation":"major_revision","confidential_remarks":"The benchmark is a useful contribution and the manuscript is well organized, but the central fair-comparison guarantee needs to be either tightened or reworded, and a baseline evaluation is needed to validate the task definitions and splits. I do not see an irreparable flaw: the test-isolation issue can be addressed by adding a hidden holdout/leaderboard or by clearly repositioning the policy as a compliance contract, and the missing baselines are a matter of additional experiments. On that basis I support major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. This paper correctly identifies the evaluation mess in channel foundation models and builds a structured benchmark: six diverse channel data sources, unit-level partitions, six task groups, a task-support matrix that disables unsupported tasks, and a careful metric contract. The curation is thoughtful—they use trajectories, sessions, flights, and vehicle links as independent units, which addresses the frame-level leakage that has inflated many published numbers. It is a solid piece of infrastructure work. But two things are missing: no baseline experiments and no actual release artifacts in the manuscript. And the central fairness guarantee—no pretraining on test units—rests on voluntary disclosure, which the paper itself concedes in Sec. VII.\n\nWhat is new: the combination of multi-domain, multi-task evaluation under one protocol, with explicit leakage-resistant partitions and a strict separation between pretraining, fine-tuning, validation, and test. The task-support matrix is a genuinely good idea: tasks that are semantically unsupported are disabled rather than given synthetic labels. The metric definitions are precise, with per-domain scoring, and the caveats about heterogeneous codebooks and non-exchangeable domains are honest.\n\nThe soft spots. The big one is the gap between the claim that any model touching a test unit cannot be presented as a compliant result and the enforcement mechanism, which is a mandatory disclosure statement. All test units are public. A participant who silently pretrains on them will beat honest submissions, and the benchmark cannot detect it. The paper acknowledges this, which is fair, but it means the benchmark's core promise of a trustworthy ranking is not yet operational. That is a real limitation, not a nitpick, though it is shared by many public benchmarks and the paper points to a hidden leaderboard as future work. The second issue is the lack of any baseline experiments. With 157,900 frames and six tasks, we get no reference numbers. Without a sanity check, you do not know whether the splits are tractable, whether the tasks are separable, or whether the leakage-resistant partitions change results at all. For a benchmark paper, that is a substantial omission. Third, the manuscript contains no data or code links. For a resource release, that is a practical blocker; the paper cannot be used as-is.\n\nWho this is for: researchers building or evaluating channel foundation models, and benchmark designers. The protocol and task definitions are worth careful study even if the release is incomplete. It deserves serious peer review, but only with conditions: baseline experiments, working links, and either a hidden test set or a reframing of the test-isolation claim as a policy dependent on community honesty. If the authors deliver those, this could become a standard evaluation tool. As it stands, it is a strong proposal waiting for evidence.","headline":"A coherent, genuinely useful benchmark protocol for CFMs, but it ships without baseline experiments or usable links, and its leakage guarantee is honor-system only.","tokens_in":13821,"tokens_out":2466,"would_cite":true,"duration_ms":27408,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CFM-Bench provides a common substrate for comparing channel foundation models across six radio configurations and six task groups, using leakage-resistant partitions and a test-exposure policy that make transfer comparisons trustworthy.","keywords":["channel foundation models","benchmark","test isolation","CSI feedback","beam prediction","localization","multi-task learning","wireless AI"],"falsifier":"Compute the average complex-CSI similarity between official training and test units and compare it with the similarity within the training set. If the cross-split similarity distribution substantially overlaps the within-split distribution, the unit-level partitioning has not removed information leakage, and rankings built on the benchmark would be inflated.","tokens_in":13051,"feed_emoji":"📶","tokens_out":6089,"duration_ms":62681,"temperature":0.7,"pith_summary":"The paper argues that current evaluations of channel foundation models (CFMs) are not comparable: each study uses its own data, splits, and metrics, so reported pretraining gains cannot be ranked across models. To fix this, the authors release CFM-Bench, which curates one representative configuration from each of six radio data sources—statistical, ray-traced, measured, and multimodal—and imposes a common evaluation contract. The contract makes test units untouchable during development, requires disclosure of all pretraining data, and defines six task groups across physical-layer, network-decision, and sensing applications. If adopted, the benchmark would let any CFM be compared fairly against other CFMs and against task-specific models, and would expose which transfer gains are real rather than artifacts of leakage or pipeline differences.","feed_headline":"One benchmark ranks channel models fairly across six radio settings","feed_subtitle":"Unit-level splits and a test-isolation policy put every pretrained channel model on equal footing.","key_machinery":"The load-bearing mechanism is unit-level leakage-resistant partitioning combined with a mandatory data-exposure policy and a task-support matrix. Partitions are drawn at the level of complete physical units so that spatially or temporally correlated samples never straddle the train/test boundary; the policy reserves official splits exclusively for fine-tuning and scoring; and the task-support matrix encodes which tasks are physically meaningful per domain, preventing superficially similar labels from being compared under incompatible semantics.","core_discovery":"The central claim is that CFM-Bench makes cross-model comparison meaningful by fixing the things that currently vary between papers. It selects one fixed radio configuration per source, partitions at the largest independent physical unit (complete trajectories, measurement sessions, vehicle links, simulation realizations, or buffered spatial regions), and forbids any benchmark split from being used in foundation-model pretraining. It also requires a data-exposure statement listing every dataset used during development, and disables scientifically unsupported task–domain combinations rather than manufacturing labels. The result is a shared substrate on which a pretrained channel representatio","pith_inferences":["The benchmark's design suggests a natural next step: adding a hidden test-set tier would close the acknowledged gap that public test sets cannot prevent repeated manual adaptation.","The task-support matrix—disabling unsupported domain–task combinations—could become a template for other foundation-model benchmarks where physical semantics vary by domain.","Because domains differ in difficulty and sample count, the macro-average score should be read with caution; per-domain inspection will likely be more informative than any single number.","The strict exclusion of tasks like temporal extrapolation on measured domains may understate what sophisticated signal processing can extract; future releases could add processed variants as separate tasks."],"forward_implications":["Any pretrained channel model can be ranked against other CFMs and against task-specific networks under identical data, splits, and metrics.","Reported pretraining gains can be checked for authenticity: gains that vanish under unit-level isolation are exposed as leakage artifacts.","Transferability can be assessed across statistical, ray-traced, measured, and multimodal channels within a single protocol.","Per-domain scores become the unit of comparison, with an unweighted macro-average explicitly demoted to a secondary summary.","Researchers get a fixed test-exposure policy that disambiguates compliant results from test-exposed or transductive ones."],"fun_headline_variants":["New benchmark fairly ranks channel foundation models across six settings","CFM-Bench unifies channel model tests, bans pretraining on splits","Test isolation lets channel foundation models compete on equal terms","Benchmark ends unfair comparisons of channel foundation models","CFM-Bench sets new standard for fair channel model evaluation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The fairness guarantee rests on voluntary disclosure and public test sets; if a participant silently uses test units during development, the benchmark's central promise of trustworthy comparison collapses.","fun_headline_variants_meta":{"raw":{"variants":["New benchmark fairly ranks channel foundation models across six settings","CFM-Bench unifies channel model tests, bans pretraining on splits","Test isolation lets channel foundation models compete on equal terms","Benchmark ends unfair comparisons of channel foundation models","CFM-Bench sets new standard for fair channel model evaluation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000462,"raw_usage":{"total_tokens":2173,"prompt_tokens":795,"completion_tokens":1378,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":1306}},"tokens_in":539,"tokens_out":1378,"duration_ms":10947,"temperature":1.0,"reasoning_tokens":1306,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:31:23.952147+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the average complex-CSI similarity between official training and test units and compare it with the similarity within the training set. If the cross-split similarity distribution substantially overlaps the within-split distribution, the unit-level partitioning has not removed information leakage, and rankings built on the benchmark would be inflated.","supporting_citations":[],"review_version":1}