{"id":"49b9aed5-15e0-48ac-9b6a-371f7c8f2af0","arxiv_id":"2508.04223","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper introduces WS-DC, a spectral-efficiency-aware codebook design that uses Wasserstein distance to align codebook activation with the optimal channel input distribution, improving inference accuracy and codebook efficiency.","lead":"This paper proposes a codebook design method for task-oriented semantic communications that explicitly optimizes how often each codeword is used, aiming to boost spectral efficiency. It introduces a Wasserstein-distance regularizer to align the learned codebook activation pattern with the optimal channel input distribution, and reports accuracy and efficiency gains.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The provided full text is a different paper (on quantum-HPC resource allocation); the abstract's claims about WS-DC and its experiments are therefore entirely unsupported by the manuscript body.","rationale":"The reader correctly identified a plausible weak assumption: the compatibility of Wasserstein regularization with task-relevant information preservation is not justified. However, the more immediate and load-bearing concern is that the manuscript body does not correspond to the abstract. The central claim is an empirical one, and the evidence for it is entirely missing. This is not an artifact of the review pipeline; it is a property of the submitted text. The reader's UNVERDICTED verdict is appropriate: there is insufficient information to validate or refute the central claim. I partially agree with the reader because the compatibility assumption is indeed a weak point, but the missing full text is the overriding issue. I do not find grounds to change the verdict, hence UNCHANGED.","tokens_in":2038,"tokens_out":2539,"duration_ms":31372,"concrete_test":"Retrieve the actual full text of arXiv:2508.04223 from arXiv. Verify that the manuscript body (Sections III–IV) defines the optimal channel input distribution and the WS-DC algorithm, and that the experimental section reports results comparing WS-DC to baseline ToSC methods on the same datasets and metrics. If the full text is indeed the quantum-HPC paper, the submission is mislinked and the central claim cannot be evaluated. As a secondary check, independently derive the gradient of the WS regularization term with respect to the encoder parameters on a simple semantic task to test whether it conflicts with task loss minimization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that WS-DC outperforms existing approaches in inference accuracy and significantly improves codebook efficiency, supported by 'experimental results.' However, the full text provided in this manuscript is Rocco et al., 'Dynamic Solutions for Hybrid Quantum-HPC Resource Allocation' (arXiv:2508.04217), which contains no mention of semantic communications, codebooks, Wasserstein distance, or any ToSC experiments. Thus every load-bearing component of the claim—the method's architecture, the definition of the 'optimal channel input distribution,' the compatibility of the WS regularizer with the task loss, and the empirical comparisons—is absent. This is not a subtle theoretical flaw but a complete lack of supporting evidence in the submitted manuscript. Even taken on the abstract alone, the assumption that steering the learned activation distribution toward an unspecified 'optimal channel input distribution' does not degrade task-relevant information is unstated and unproven. The abstract asserts both task accuracy and spectral efficiency gains without presenting the mechanism or the measurements. In its current form, the manuscript provides no basis for evaluating the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract proposes a spectral-efficiency-aware codebook design framework, WS-DC, for task-oriented semantic communications. The claimed method combines a task loss with a Wasserstein-distance regularizer that steers the learned codebook activation distribution toward an 'optimal channel input distribution', with experimental results said to show improved inference accuracy and codebook efficiency. However, the full text of the submitted manuscript is an entirely different paper on dynamic solutions for hybrid quantum-HPC resource allocation (arXiv:2508.04217). It contains no mention of semantic communications, codebooks, Wasserstein distance, the WS-DC scheme, or any experiment related to the abstract's claims. As submitted, the manuscript therefore provides no method description, no definitions, no theoretical development, and no empirical evidence for its central claims.","tokens_in":2314,"tokens_out":1712,"duration_ms":21168,"significance":"If the proposed WS-DC framework were fully developed and validated, it could be a useful contribution to codebook design for task-oriented semantic communications, particularly in addressing sparse codebook activation and spectral efficiency. However, because the submitted manuscript body is a different paper entirely, none of the claimed contributions can be evaluated. The significance of the work, assuming it exists, cannot be assessed from the submitted text.","major_comments":[{"comment":"The central claims of the abstract are entirely unsupported by the manuscript body. The full text is 'Dynamic Solutions for Hybrid Quantum-HPC Resource Allocation' by Rocco et al., which discusses malleable job scheduling in HPC-QC systems and contains no reference to semantic communications, codebooks, Wasserstein distance, spectral efficiency, or WS-DC. Every load-bearing component of the claimed contribution—the architecture, the objective function, the definition of 'optimal channel input distribution', the generative reinterpretation of WS theory, and the experimental comparisons—is absent. As submitted, the paper cannot be reviewed for its stated content.","section":"Abstract vs. Full Text"},{"comment":"Even taken on its own, the abstract does not define the 'optimal channel input distribution' that is used as the WS regularization target, nor does it explain how minimizing the Wasserstein distance to that target is compatible with preserving task-relevant information. The claim that WS-DC 'outperforms existing approaches in inference accuracy and significantly improves codebook efficiency' is presented without any of the necessary experimental details: datasets, baselines, channel models, communication protocols, metrics, or error bars. These omissions make the abstract's empirical claims unverifiable.","section":"Abstract"},{"comment":"The full text's introduction and subsequent sections describe a resource-allocation problem for hybrid quantum-HPC systems, including references to QPU scheduling, malleable jobs, and classical-quantum integration. None of this material connects to the abstract's task-oriented semantic communication framework. The manuscript therefore does not contain the proposed method or any of its supporting derivations, making the central claim of the paper unsupported by the submitted evidence.","section":"Full Text, Section I"}],"minor_comments":[{"comment":"The manuscript appears to contain the wrong full text relative to its abstract. At minimum, the title, author list, and abstract should correspond to the body. This is not a stylistic issue but a fundamental mismatch.","section":"General"},{"comment":"The phrase 'optimal channel input distribution' is used without a formal definition or reference. If this distribution is meant to be capacity-achieving, that should be stated explicitly; if it is learned, its role in the WS regularizer needs clarification.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"This submission appears to be a file mismatch: the full text is a different arXiv paper (2508.04217) on quantum-HPC resource allocation. As submitted, there is no content to review for the claimed topic. The editor may wish to verify whether an upload error occurred, but under the current submission the manuscript cannot be accepted or meaningfully revised without complete replacement of the text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You need to know two things. First, the full text of arXiv:2508.04223 is not the paper advertised. The title and abstract are about semantic communication codebooks; the body is Rocco et al., \"Dynamic Solutions for Hybrid Quantum-HPC Resource Allocation.\" That is not a typo in a section; it's the entire manuscript. So there is nothing I can actually review beyond the abstract.\n\nSecond, the abstract itself points at a legitimate problem and a plausible fix. Sparse activation of learned codebooks wastes channel capacity in task-oriented semantic communications. Adding a Wasserstein-distance penalty to pull the empirical activation distribution toward the channel-optimal input distribution is a sensible thing to try, and I don't recall seeing that exact regularization in the ToSC literature. The \"generative perspective\" language is vague but harmless.\n\nNow the soft spots, in proportion. With the abstract alone, the central claim—better accuracy and better spectral efficiency—is unsupported. There is no definition of the \"optimal channel input distribution.\" If that target is the capacity-achieving distribution, nothing in the abstract says the penalty won't fight the task loss. The compatibility assumption is unstated and unproven. No datasets, baselines, or error bars are given. And the actual manuscript text, being the wrong paper, contains none of the method or experiments. So I can't comment on the math, the data, or the citation pattern; there is no basis.\n\nIf the correct manuscript exists, this could be a reasonable incremental contribution, not a breakthrough. But as submitted, it is not reviewable. My recommendation to the editor: send it back and ask for the correct full text before anything else. That is a desk-reject-as-is, not a scientific judgment about the underlying idea.\n\nAll best.","headline":"The full manuscript text is a different paper, so the abstract's claims are unauditable—this needs to go back to the authors before any review.","tokens_in":2730,"tokens_out":2769,"would_cite":false,"duration_ms":30924,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Task-oriented semantic codebooks can be made both more accurate and more spectrally efficient by adding a Wasserstein-distance regularizer that pulls the codebook activation distribution toward the channel-optimal input distribution.","keywords":["task-oriented semantic communications","codebook design","spectral efficiency","Wasserstein distance","activation distribution","channel-aware latent representations","digital semantic communications","capacity-approaching"],"falsifier":"Train the same semantic encoder with and without the WS regularizer, logging activation frequencies and end-to-end inference accuracy across SNR from 0 to 20 dB. If every nonzero regularization weight makes accuracy worse than the unregularized baseline, or if the activation distribution's Wasserstein distance to a properly defined channel-optimal input distribution does not decrease when accuracy improves, the proposed mechanism is not what produces the reported gains.","tokens_in":1977,"feed_emoji":"📡","tokens_out":7699,"duration_ms":80749,"temperature":0.7,"pith_summary":"This paper tackles a hidden waste in task-oriented semantic communication: learned codebooks that map semantic features to channel symbols are often sparsely activated, so most symbols carry no useful task information. The authors build a codebook design framework that makes activation probability part of the optimization objective, so a codebook is judged by both inference accuracy and how much of the channel capacity it actually uses. They regularize the learned activation distribution toward the optimal channel input distribution with Wasserstein distance, the minimal transport cost between two distributions, and reinterpret that distance from a generative angle to fit semantic communication. Their proposed scheme, WS-DC, is reported in the abstract to outperform existing approaches in inference accuracy while significantly improving codebook efficiency. If that holds, semantic links can approach channel capacity without sacrificing the task they serve.","feed_headline":"Semantic codebook with WS regularizer wins on accuracy and spectrum","feed_subtitle":"By steering activation distribution toward channel-optimal inputs, WS-DC packs more task-relevant bits per symbol.","key_machinery":"The load-bearing component is a training objective in which the codebook activation probability enters explicitly, regularized by the Wasserstein distance between the learned activation distribution and the optimal channel input distribution. WS-DC (the Wasserstein-based adaptive hybrid distribution scheme) combines the task loss with this WS regularizer: the hybrid distribution is the codebook's activation distribution, and the Wasserstein term supplies channel-awareness. This machinery converts spectral efficiency from an incidental byproduct of quantization into an explicit, differentiable design target.","core_discovery":"The paper claims that the sparse activation of learned codebooks is the main spectral-efficiency bottleneck in digital task-oriented semantic communication, and that this bottleneck can be optimized explicitly. The proposed WS-DC scheme learns compact, task-driven, channel-aware latent representations by adding a Wasserstein-distance term to the training loss; this term pushes the distribution of codebook activations toward the optimal channel input distribution, i.e., the input distribution that is best matched to the channel. The authors also reinterpret Wasserstein distance from a generative perspective, treating the codebook less as a fixed quantizer and more as a generator of channel in","pith_inferences":["Editorial note: the full-text body supplied under this title is a different manuscript, on hybrid quantum-HPC resource allocation, and contains no derivations or experimental detail for the abstract's claims; the extraction above therefore rests on the abstract alone.","The central compatibility assumption is testable: an explicit characterization of the optimal channel input distribution for each channel model would let one check whether WS-DC really approaches it, or whether the gains come from merely balancing codebook usage.","The reweighted objective resembles rate-distortion-perception tradeoffs, suggesting WS-DC could be viewed as a semantic analogue in which the perception constraint is replaced by a channel-capacity constraint on the codebook's output distribution."],"forward_implications":["Task-oriented semantic systems using WS-DC should transmit fewer channel symbols for the same inference task, since rarely used codebook entries no longer consume capacity.","The channel-awareness added by the Wasserstein regularizer should make codebook performance transfer across signal-to-noise conditions, rather than being optimized for one fixed channel.","Reinterpreting Wasserstein distance in generative terms opens codebook design to generative-model tools, framing the codebook as a learned generator of near-optimal channel inputs.","The framework makes the tradeoff between inference accuracy and spectral efficiency explicit and tunable through the weight of the WS regularizer.","During deployment, the same activation-probability logging used in training gives an online diagnostic of how far the current codebook usage is from channel-optimal usage."],"supporting_citations":[],"fun_headline_variants":["WS-DC codebook: accuracy up, spectrum cost down","Wasserstein regularizer steers codebook to channel-optimal usage","Spectral-efficiency codebook: WS-DC beats prior semantic comms","Align activation distribution with channel capacity via Wasserstein","Task-driven codebook with WS regularizer: less spectrum, more accuracy"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The claim stands or falls on whether pulling the codebook activation distribution toward the optimal channel input distribution can be done without erasing the differences between codewords that the inference task depends on.","fun_headline_variants_meta":{"raw":{"variants":["WS-DC codebook: accuracy up, spectrum cost down","Wasserstein regularizer steers codebook to channel-optimal usage","Spectral-efficiency codebook: WS-DC beats prior semantic comms","Align activation distribution with channel capacity via Wasserstein","Task-driven codebook with WS regularizer: less spectrum, more accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000167,"raw_usage":{"total_tokens":1087,"prompt_tokens":733,"completion_tokens":354,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":278}},"tokens_in":477,"tokens_out":354,"duration_ms":5469,"temperature":1.0,"reasoning_tokens":278,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:45:26.063861+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same semantic encoder with and without the WS regularizer, logging activation frequencies and end-to-end inference accuracy across SNR from 0 to 20 dB. If every nonzero regularization weight makes accuracy worse than the unregularized baseline, or if the activation distribution's Wasserstein distance to a properly defined channel-optimal input distribution does not decrease when accuracy improves, the proposed mechanism is not what produces the reported gains.","supporting_citations":[],"review_version":1}