{"id":"34631a31-79bf-456c-afba-a58b58e0a05d","arxiv_id":"2606.31365","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Extends probing evaluation to acoustic teleportation codecs and reports that speaker identity stays partitioned while room acoustics leak into speech embeddings, with acoustic embeddings estimating room parameters nearly as well as supervised baselines.","lead":"The paper extends a probing framework to test disentanglement in neural audio codecs by checking how well room acoustics and speaker identity can be recovered from different latent partitions. A smart generalist might read it to see a practical way to verify whether AI audio tools truly keep separate the content, voice, and environment without unwanted mixing.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Probing gap may not isolate true disentanglement if probes exploit data correlations or indirect routes","rationale":"The reader's weakest assumption directly identifies the load-bearing point. Full-text details on probe architectures, training data statistics, and any correlation controls would be needed to close the gap; absent those, the probing metric remains the primary uncertainty and the UNVERDICTED stance is appropriate.","tokens_in":1650,"tokens_out":297,"duration_ms":28014,"concrete_test":"Construct a controlled test set where speaker identity and room parameters are statistically independent (by resampling or synthetic mixing); retrain or re-evaluate the same probes; if the gap on the speech partition for acoustics shrinks substantially, the original leakage conclusion is driven by dataset correlations rather than the training objective.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that speaker identity is confined while acoustics leak into speech embeddings—rests on the performance gap between probes on intended vs. unintended partitions. This gap is taken as evidence of leakage 'due to the training objective.' However, if speaker and room parameters are correlated in the training data, or if the codec's latent space allows information to be accessed via non-partition routes, the unintended-partition probe could succeed without actual leakage into the speech embeddings. The method does not appear to include controls that would rule this out (e.g., decorrelated test sets or information-theoretic bounds).","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes extending a probing-based evaluation framework for disentanglement in acoustic teleportation neural audio codecs, moving beyond cross-reconstruction quality metrics. Probes regress room-acoustic parameters (reverberation time, clarity, direct-to-reverberant ratio) and classify speaker identity on latent partitions; the performance gap between intended and unintended partitions quantifies leakage. Applied to one such codec, the results indicate speaker identity is largely confined to its partition while acoustics leak into speech embeddings due to the training objective, and that acoustic embeddings can estimate room parameters within 0.02 s of supervised baselines, suggesting emergent physically meaningful structure without explicit supervision.","tokens_in":1776,"tokens_out":522,"duration_ms":46729,"significance":"If the probing gaps can be shown to isolate true disentanglement, the work supplies a more direct assessment tool than cross-reconstruction for models used in voice conversion and acoustic teleportation. The reported near-parity of unsupervised acoustic embeddings with supervised room-parameter estimators is a concrete, falsifiable observation that could guide future codec design. The paper explicitly contrasts its approach with prior reliance on reconstruction quality, which is a useful contribution to evaluation methodology in the field.","major_comments":[{"comment":"Abstract: the quantitative claim that acoustic embeddings estimate room parameters 'within 0.02 s of supervised baselines' is presented without any description of probe architectures, training regimes, data splits, or statistical controls; this detail is load-bearing for the claim that physically meaningful structure emerges without explicit supervision.","section":"Abstract"},{"comment":"Evaluation section (probing framework): the central claim that the intended-vs-unintended partition performance gap measures leakage 'due to the training objective' rests on the assumption that probes cannot exploit data correlations between speaker and room parameters or indirect latent routes; the manuscript provides no decorrelated test sets, information-theoretic bounds, or ablation controls to rule these out, directly affecting the validity of the disentanglement conclusions.","section":"Evaluation section (probing framework)"}],"minor_comments":[{"comment":"Abstract: the abbreviation 's' in '0.02 s gap' is ambiguous (seconds?); expand or define the metric explicitly.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The absence of methodological detail even in the abstract suggests the manuscript may still be at a preliminary stage; the journal's standards for empirical audio-signal-processing papers typically require full reproducibility information for quantitative claims of this type."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments. Below we address each major comment point by point.","responses":[{"response":"The probe architectures (linear regressors), training regimes, data splits, and controls are described in the Evaluation section. We agree the abstract presents the claim without sufficient context. We will revise the abstract to add a brief qualifier referencing the linear probes on held-out data.","revision_made":"partial","referee_comment":"[Abstract] Abstract: the quantitative claim that acoustic embeddings estimate room parameters 'within 0.02 s of supervised baselines' is presented without any description of probe architectures, training regimes, data splits, or statistical controls; this detail is load-bearing for the claim that physically meaningful structure emerges without explicit supervision."},{"response":"We acknowledge this limitation: the manuscript uses standard dataset splits without decorrelated test sets, information-theoretic bounds, or ablations for indirect routes or speaker-room correlations. We will add a dedicated limitations paragraph qualifying the interpretation of leakage as due to the training objective and noting the assumption. The asymmetric performance gaps still provide supporting evidence, but we will tone down causal language.","revision_made":"partial","referee_comment":"[Evaluation section (probing framework)] Evaluation section (probing framework): the central claim that the intended-vs-unintended partition performance gap measures leakage 'due to the training objective' rests on the assumption that probes cannot exploit data correlations between speaker and room parameters or indirect latent routes; the manuscript provides no decorrelated test sets, information-theoretic bounds, or ablation controls to rule these out, directly affecting the validity of the disentanglement conclusions."}],"tokens_in":1347,"tokens_out":361,"duration_ms":45500,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper takes an existing probing idea and applies it to room-acoustic parameters and speaker identity on acoustic teleportation codecs. It reports that speaker information stays mostly in its own partition while acoustics leak into the speech embeddings, and that the acoustic partition predicts reverberation time, clarity, and DRR within 0.02 s of a supervised baseline.\n\nWhat is new is the targeted use of these room parameters as probe targets for this codec type. Cross-reconstruction alone cannot catch the leakage the authors describe, so adding regression and classification probes on the partitions is a direct way to look for it. The observation that the acoustic embeddings carry usable room information without explicit supervision is also worth noting if the numbers hold.\n\nThe soft spot is the missing controls. The abstract gives no probe architecture details, training splits, or tests for whether speaker and room parameters are correlated in the data. If they are, a probe on the unintended partition could succeed without any actual leakage into the speech embeddings. The stress-test concern lands here: the performance gap is treated as evidence of leakage due to the training objective, but without decorrelated test sets or information bounds the gap could come from other routes. That makes the central claim harder to assess from what is shown.\n\nThis is for people already working on disentangled neural audio codecs who need evaluation tools beyond reconstruction metrics. It is worth sending to peer review so the methods section can be checked for those controls and for whether the probe results survive them. If the controls are absent or weak, the leakage conclusion would need to be walked back.","headline":"Probing extension for codec disentanglement is a sensible step but the leakage claim rests on a gap that could reflect data correlations instead.","tokens_in":2259,"tokens_out":388,"would_cite":false,"duration_ms":33665,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Probing reveals speaker identity stays mostly in its partition while acoustics leak into speech embeddings in neural codecs.","keywords":["disentanglement evaluation","neural audio codecs","acoustic teleportation","probing methods","room acoustics","speaker identity","voice conversion","latent representations"],"falsifier":"If acoustic embeddings estimate room parameters no better than chance levels or if speaker classification accuracy shows no gap between intended and unintended partitions under varied probe models, the claims on confinement and leakage would not hold.","tokens_in":2552,"feed_emoji":"🎙️","tokens_out":524,"duration_ms":40705,"temperature":0.7,"pith_summary":"The paper introduces a probing method to evaluate how well neural audio codecs separate speech content, speaker identity, and room acoustics into distinct latent spaces. Probes regress room parameters such as reverberation time and classify speakers from each partition, with the performance difference between intended and unintended partitions serving as the disentanglement score. This approach finds that speaker identity largely remains in its assigned space but acoustic details leak into speech embeddings because of the codec's training goal. Acoustic embeddings estimate room parameters nearly as well as models trained with labels, showing that physical structure can appear without direct supervision. The method addresses the limits of cross-reconstruction tests that miss such leakage.","feed_headline":"Probes show acoustics leak into speech embeddings in codecs","feed_subtitle":"Speaker identity stays mostly partitioned but room parameters emerge unsupervised in acoustic embeddings and leak elsewhere.","key_machinery":"Probing framework that regresses room-acoustic parameters and classifies speaker identity from latent partitions, using performance gaps between intended and unintended partitions as the disentanglement measure.","core_discovery":"The probing-based evaluation demonstrates that in an acoustic teleportation codec, speaker identity is largely confined to its intended partition while acoustic information leaks into the speech embeddings due to the training objective. Acoustic embeddings estimate room parameters within 0.02 seconds of supervised baselines without explicit supervision.","pith_inferences":["This probing approach could test disentanglement in other latent variable models for audio or images.","Reducing the observed leakage might improve performance in downstream tasks like voice conversion.","The unsupervised emergence of room parameter estimates suggests latent spaces can implicitly learn physical acoustic models."],"forward_implications":["Speaker embeddings primarily hold identity information with little acoustic leakage.","Acoustic embeddings capture meaningful room properties such as reverberation time and clarity without explicit labels.","Cross-reconstruction quality alone fails to detect leakage across partitions.","Training objectives must be adjusted to reduce acoustic information in speech embeddings for cleaner separation."],"fun_headline_variants":["Probes expose acoustics leaking into speech embeddings","Speaker identity stays confined to its codec partition","Acoustic embeddings estimate room params unsupervised","Probing detects leakage across codec partitions"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The performance gap between probes on intended versus unintended partitions provides a reliable and unbiased measure of disentanglement without the probes themselves suffering from leakage or selection effects.","fun_headline_variants_meta":{"raw":{"variants":["Probes expose acoustics leaking into speech embeddings","Speaker identity stays confined to its codec partition","Acoustic embeddings estimate room params unsupervised","Probing detects leakage across codec partitions"]},"model":"grok-4.3","cost_usd":0.0043,"raw_usage":{"total_tokens":2106,"prompt_tokens":557,"num_sources_used":0,"completion_tokens":51,"cost_in_usd_ticks":42999500,"prompt_tokens_details":{"text_tokens":557,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1498,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":557,"tokens_out":51,"duration_ms":18690,"temperature":1.0,"reasoning_tokens":1498,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T03:43:26.568739+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If acoustic embeddings estimate room parameters no better than chance levels or if speaker classification accuracy shows no gap between intended and unintended partitions under varied probe models, the claims on confinement and leakage would not hold.","supporting_citations":[],"review_version":1}