{"id":"d4f43f5d-9c7a-414f-afe3-be5224715000","arxiv_id":"2508.02849","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"SecoustiCodec achieves state-of-the-art low-bitrate speech coding by disentangling semantic and paralinguistic features into a single codebook.","lead":"This paper proposes SecoustiCodec, a streaming speech codec that compresses speech to very low bitrates while separating meaning from voice style. It claims state-of-the-art audio quality (PESQ 2.58 at 1 kbps) and plans to release code and weights.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SOTA PESQ claim is not checkable from the available text; a direct reproduction from released code and weights is required before the claim can be accepted.","rationale":"The reader's verdict of UNVERDICTED is correct, and the rationale that the full text is unusable is confirmed by the corrupted physics header. The reader's weakest_assumption focuses on the contrastive alignment mechanism; that is the most natural technical weak point in the abstract, but it is not the most immediate blocker. The most load-bearing issue is that the central quantitative claim, SOTA PESQ at extreme low bitrates, is presented without any accessible methodology, dataset definition, or baseline comparison. A PESQ number without the evaluation configuration is not evidence. Agreement is partial because the reader identifies a plausible technical fragility, whereas my stress-test emphasizes that the claim cannot even be tested from the supplied manuscript. The verdict should remain UNCHANGED because no new technical objection can be responsibly raised on the basis of an abstract alone; the paper is simply unverified. The recommended concrete test is a faithful reproduction of the headline numbers from the released code and weights, with the speaker-classifier probe as a targeted check of the disentanglement claim should the full paper become available.","tokens_in":1008,"tokens_out":3361,"duration_ms":43317,"concrete_test":"Retrieve the actual SecoustiCodec PDF from arXiv:2508.02849 and rerun the reported 0.27 kbps and 1 kbps evaluations using the released code and model weights on the same evaluation set and against the same baselines as the paper's comparison table. If the PESQ values reproduce within 0.1 and the baseline protocol is matched, the headline claim is supported; if the artifacts are absent or the numbers do not reproduce, the claim should remain unverified. As a secondary probe of the disentanglement assumption, train a speaker or emotion classifier on the output semantic token sequence; high accuracy would indicate residual paralinguistic leakage and would contradict the claim that contrastive alignment removes paralinguistic information.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that SecoustiCodec achieves SOTA reconstruction quality with PESQ 1.77 at 0.27 kbps and 2.58 at 1 kbps. This is a quantitative comparison claim, but the available full text is not the paper: it is a corrupted arXiv header for an unrelated physics.optics manuscript. No method section, evaluation protocol, dataset list, baseline table, or ablation survives. Consequently, the PESQ numbers are unsupported assertions in the current submission: there is no way to check what was measured, on which utterances, against which codecs, or with what configuration. This is not a theoretical dispute; it is a completeness and support problem. Even the reader's technical concern about the contrastive disentanglement assumption cannot be evaluated because the training objective, model architecture, and any supporting figures or equations are absent. The abstract also contains a self-contradiction about artifact availability, stating both that code and model weights 'will be open-sourced upon the completion of the peer-review process' and that they have already been open-sourced; this does not affect the PESQ claim directly, but it further undermines reproducibility. Given the available evidence, the SOTA reconstruction claim is unverified, and the appropriate disposition is to withhold judgment until the actual manuscript and artifacts are inspectable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This submission proposes SecoustiCodec, a low-bitrate streaming speech codec with a single codebook, claiming cross-modal alignment between text and speech, VAE-FSQ semantic quantization, paralinguistic encoding, and multi-stage optimization. The abstract reports state-of-the-art PESQ scores of 1.77 at 0.27 kbps and 2.58 at 1 kbps. However, the full text provided for review is not the SecoustiCodec manuscript; it is garbled text from an unrelated physics.optics paper (EPFL LNET, E-mail giulia.tagliabue@epfl.ch, arXiv:2508.02850v1). No architecture details, training objectives, evaluation protocol, baselines, results, or ablation studies are present in the submission.","tokens_in":1233,"tokens_out":3107,"duration_ms":31791,"significance":"If the claimed results were substantiated, SecoustiCodec would offer a meaningful advance: a single-codebook codec supporting streaming at ultralow bitrates with disentangled semantic and paralinguistic information could benefit speech-text language models. The significance cannot be assessed, however, because the manuscript body is absent. A quantitative SOTA claim without any experimental section is not a contribution in the current form.","major_comments":[{"comment":"The submission's full text is an unrelated physics.optics manuscript (EPFL LNET, arXiv:2508.02850v1) rather than the SecoustiCodec paper. All claimed technical content—the contrastive learning alignment, VAE-FSQ quantization, paralinguistic encoding, multi-stage optimization, and streaming architecture—is therefore absent. This is a load-bearing completeness failure: the reviewer cannot verify the central claims.","section":"Full text"},{"comment":"The abstract states SOTA PESQ of 1.77/2.58 at 0.27/1 kbps but gives no dataset, test set, baseline codecs, training configuration, or uncertainty estimates. Even if the full text were present, a PESQ comparison without these details would not be checkable; in this submission it is a bare unsupported assertion.","section":"Abstract"},{"comment":"The artifact availability statements are contradictory: the abstract says code and weights 'will be open-sourced upon the completion of the peer-review process' and in the next sentence says 'We've open-sourced SecoustiCodec's demo, code, and model weights.' This inconsistency makes it impossible to determine the actual availability of the claimed artifacts.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract references Figure ~\\ref{fig:pesq_kbps_below_2kbps}, but no figures or body text accompany the submission, so the referenced result cannot be located.","section":"Abstract"},{"comment":"The full text is not decodable as prose; the encoding is corrupted beyond recovery, which in itself blocks review.","section":"Full text"},{"comment":"No references are provided in the submission, so related-work context and prior codec comparisons are entirely missing.","section":"Full text"}],"recommendation":"reject","confidential_remarks":"This is not a content review; the uploaded manuscript is the wrong file. The arXiv identifier 2508.02849 is listed as eess.AS, but the body text is from a physics.optics preprint. I recommend rejecting the current submission and inviting the authors to resubmit the correct PDF, at which point a substantive review can begin. The contradictory open-source statements should also be resolved before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The submission is not reviewable in its current form: the full text is a corrupted physics.optics manuscript, so only the abstract is usable. On the abstract alone, SecoustiCodec looks like a plausible step forward in low-bitrate streaming codecs—the problem it targets (residual paralinguistic info in semantic tokens, poor semantic completeness, no streaming support) is real, and the proposed combination of VAE+FSQ for semantic-only quantization, contrastive text-speech alignment to strip paralinguistic content, and a single-codebook streaming design is a coherent architectural idea. The abstract states these components clearly, and the reported PESQ figures (1.77 at 0.27 kbps, 2.58 at 1 kbps) are not implausible for the subfield.\n\nThat is where the credit ends. The abstract provides no experimental protocol: no baselines, datasets, error bars, or ablation. The central claim—that contrastive learning aligns text and speech in a joint frame-level space and thereby removes paralinguistic information without hurting reconstruction—is load-bearing but cannot be evaluated because the training objective and architecture are not shown. Even the artifact statement is self-contradictory: one line says code and weights will be open-sourced after peer review, the next says they are already open-sourced. Minor, but it does not inspire confidence.\n\nThe stress-test note is right: this is a completeness and support problem, not a theoretical dispute. There is no way to check what was measured, on which utterances, against which codecs, or with what configuration. I cannot see circularity or invented entities, but with zero experimental detail I also cannot rule out evaluation-set tuning. The manuscript as submitted is incomplete, and the appropriate disposition is to withhold judgment.\n\nRecommendation: desk reject this submission for incompleteness. If the authors resubmit the actual paper with code and weights, it deserves a serious referee. As it stands, I would not cite it or bring it to reading group.","headline":"The abstract describes an interesting low-bitrate streaming codec, but the submitted full text is an unrelated physics paper, so none of the claims can be checked.","tokens_in":1792,"tokens_out":1845,"would_cite":false,"duration_ms":22390,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SecoustiCodec claims state-of-the-art speech reconstruction at 0.27 kbps and 1 kbps by disentangling semantic from paralinguistic information in a single-codebook space.","keywords":["speech codec","low-bitrate audio","single-codebook quantization","semantic disentanglement","contrastive learning","finite scalar quantization","streaming speech codec","PESQ evaluation"],"falsifier":"Take recordings of the same sentence spoken by different speakers with different emotions and run them through the semantic encoder: if the resulting token sequences differ by more than a small tolerance, the semantic stream still carries paralinguistic information and the disentanglement claim fails. Conversely, re-synthesizing from semantic tokens plus a mismatched paralinguistic code should change voice and emotion but not the words; if the words change, the codebook has not separated the two.","tokens_in":828,"feed_emoji":"🎙️","tokens_out":5599,"duration_ms":57894,"temperature":0.7,"pith_summary":"This paper proposes SecoustiCodec, a low-bitrate streaming speech codec that compresses speech into a single codebook while keeping reconstruction quality high. The central claim is that semantic content (what was said) and paralinguistic content (timbre, emotion, prosody) can be separated inside that codebook, with a contrastive text-speech alignment removing paralinguistic information from the semantic tokens and a separate paralinguistic encoding closing the information gap. The paper reports state-of-the-art reconstruction quality, PESQ 1.77 at 0.27 kbps and 2.58 at 1 kbps, which would make the codec attractive for speech-text language models and for streaming. If the claim holds, very low bitrates no longer force a choice between semantic completeness and faithful reconstruction.","feed_headline":"Speech codec separates meaning from voice at 0.27 kbps","feed_subtitle":"SecoustiCodec's single-codebook design reports PESQ 1.77 at 0.27 kbps and 2.58 at 1 kbps.","key_machinery":"The load-bearing mechanism is the contrastive alignment of speech and text embeddings in a joint multimodal frame-level space. In that space, the model learns to keep in the semantic code only what is common to the words in both modalities, so timbre and emotion are pushed into a separate paralinguistic code; an acoustic-constrained multi-stage optimization supplies what is missing for reconstruction. The FSQ-based VAE quantizer counters the long-tail distribution of token usage, which the paper credits for high codebook utilization at low bitrate.","core_discovery":"The paper claims that a single-codebook architecture can outperform prior speech codecs at extremely low bitrates because it stops trying to make one token stream carry both meaning and voice. Instead, SecoustiCodec disentangles these two types of information: a semantic quantizer built on a variational autoencoder with finite scalar quantization encodes only linguistic content, while a paralinguistic encoder fills in the acoustic details the semantic stream drops. The disentanglement is driven by contrastive learning that aligns speech with text in a shared frame-level multimodal space, which the paper says removes timbre, emotion, and other paralinguistic traces from the semantic tokens. A multi-stage, acoustically constrained optimization keeps this joint training stable. The result reported is state-of-the-art reconstruction quality, PESQ 1.77 at 0.27 kbps and 2.58 at 1 kbps, with streaming support and a single codebook instead of a hierarchy of codebooks.","pith_inferences":["Beyond the paper's comparisons, the disentanglement claim is directly testable: feed recordings of the same sentence spoken by different speakers and emotions through the semantic encoder, and check whether the semantic token sequence remains nearly identical.","The information-gap design implies a graceful trade-off: as the paralinguistic bit budget shrinks, reconstruction quality should fall while semantic accuracy holds; measuring that curve would isolate the contribution of the paralinguistic encoder.","If the FSQ quantizer truly fixes the long-tail problem, codebook utilization should stay high even at 0.27 kbps, a statistic the paper does not report but that would let others verify the mechanism."],"forward_implications":["Speech-text language models could be driven by one semantic token stream below 1 kbps, instead of multiple stacked codebooks, simplifying the audio-text interface.","Streaming becomes practical at very low bitrates because SecoustiCodec is designed as a streaming codec rather than a whole-utterance codec.","If the disentanglement is real, editing voice, emotion, or prosody reduces to swapping the paralinguistic code while leaving the semantic token sequence fixed.","The reported numbers set a new operating point on the bitrate-quality curve: PESQ 1.77 at 0.27 kbps and 2.58 at 1 kbps."],"supporting_citations":[],"fun_headline_variants":["Single-codec speech at 0.27 kbps beats prior codecs","Streaming codec disentangles semantics at 0.27 kbps","Low-bitrate codec splits meaning from voice at 0.27 kbps","Speech codec strips voice from meaning, hits 0.27 kbps","Single-codebook codec: SOTA PESQ at 0.27 kbps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the premise that contrastive text-speech alignment genuinely separates semantic from paralinguistic content in the joint space, and that removing that paralinguistic information does not discard acoustic detail needed for reconstruction.","fun_headline_variants_meta":{"raw":{"variants":["Single-codec speech at 0.27 kbps beats prior codecs","Streaming codec disentangles semantics at 0.27 kbps","Low-bitrate codec splits meaning from voice at 0.27 kbps","Speech codec strips voice from meaning, hits 0.27 kbps","Single-codebook codec: SOTA PESQ at 0.27 kbps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000789,"raw_usage":{"total_tokens":3535,"prompt_tokens":1060,"completion_tokens":2475,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":2371}},"tokens_in":676,"tokens_out":2475,"duration_ms":18506,"temperature":1.0,"reasoning_tokens":2371,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:49:40.580643+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take recordings of the same sentence spoken by different speakers with different emotions and run them through the semantic encoder: if the resulting token sequences differ by more than a small tolerance, the semantic stream still carries paralinguistic information and the disentanglement claim fails. Conversely, re-synthesizing from semantic tokens plus a mismatched paralinguistic code should change voice and emotion but not the words; if the words change, the codebook has not separated the two.","supporting_citations":[],"review_version":1}