{"id":"58b7a8a6-5b9f-404a-b0e8-c6dca347bcd8","arxiv_id":"2508.12709","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Abstract promises SOTA audio SSL via Multiple Choice Learning; full text is a mismatched PDE estimation paper, leaving the audio claims entirely unsupported.","lead":"The submitted preprint presents an abstract about MATPAC++, a self-supervised audio representation model, but the full text is an unrelated paper on estimating PDE and delay-PDE models from data. As a result, the claimed audio state-of-the-art results have no supporting method, experiments, or tables in the manuscript.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim is unsupported because the submitted full text is a different paper on PDE estimation, not the MATPAC++ audio SSL work described in the abstract.","rationale":"The reader identified the same load-bearing issue: the submitted full text does not correspond to the claimed MATPAC++ work, so the central empirical superiority claim is unsupported. My stress test confirms this directly from the text: the body is a PDE estimation paper with no audio SSL content. The only way the central claim could be true is if the submission accidentally omitted the real MATPAC++ text, but as submitted that is not verifiable. I considered whether there might be a second concern about the MCL modeling premise itself (that explicit multi-hypothesis prediction improves representation quality), but that concern cannot be evaluated because the method is not described. Thus the primary concern is the mismatch between abstract and body. The reader's REJECT verdict with low confidence is appropriate; I do not see a reason to change it. I have not identified any independent supporting evidence in the submission, such as machine-checked proofs or reproducible code, that would mitigate this issue.","tokens_in":7459,"tokens_out":1702,"duration_ms":23591,"concrete_test":"Perform an automated keyword search over the entire submitted text (including appendices) for the strings 'MATPAC', 'AudioSet', 'Multiple Choice Learning', 'MCL', and 'masked latent'. Also extract the PDF title and author block and compare against the abstract's claimed subject. If no occurrences are found and the title/author block corresponds to the PDE paper, the central MATPAC++ claim has zero supporting content in the submission, confirming the rejection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that MATPAC++ achieves state-of-the-art results on AudioSet and downstream audio tasks through Multiple Choice Learning (MCL) integrated into MATPAC's prediction and unsupervised classification pretext tasks. For this claim to hold, the submitted document must contain the MATPAC++ architecture, the MCL formulation, training details, and evaluation results. None of these appear in the provided full text. Instead, the body is a physics paper titled 'Hyperparameter Optimization in the Estimation of PDE and Delay-PDE models from data' (matching arXiv:2508.12715), with no mention of MATPAC, AudioSet, MCL, masked latent prediction, or any audio representation learning. This is not an internal inconsistency or a debatable modeling choice; it is a complete absence of evidence for the stated claim. The reader's REJECT verdict is therefore appropriate, though the problem is more fundamental than an empirical weakness: the claim cannot be checked at all. I am not treating the unrelated PDE content as fraudulent; taken on its own it may be a legitimate manuscript, but it does not provide any support for the abstract's assertions about MATPAC++.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract of arXiv:2508.12709 announces MATPAC++, a self-supervised audio representation learning method that integrates Multiple Choice Learning (MCL) into the prediction and unsupervised classification pretext tasks of MATPAC, and claims state-of-the-art results on AudioSet fine-tuning and downstream linear-probing tasks, as well as improved efficiency for music-only training. However, the full text supplied with the submission is not the MATPAC++ paper. It is a manuscript titled 'Hyperparameter Optimization in the Estimation of PDE and Delay-PDE models from data', with its own physicists' abstract, introduction, methods, synthetic benchmark experiments on Allen-Cahn/Cahn-Hilliard and reaction-diffusion systems, and a bibliography on sparse identification of dynamical systems. The body contains no mention of MATPAC, MATPAC++, MCL, masked latent prediction, audio representation learning, AudioSet, or any of the claimed experiments. The submitted document therefore provides no derivations, no architectural description, no training details, and no evaluation results for the claims made in the abstract.","tokens_in":7771,"tokens_out":1573,"duration_ms":20960,"significance":"If the MATPAC++ results described in the abstract were properly supported, the contribution could be significant: the idea of using MCL to model predictive ambiguity in masked audio SSL is a sensible direction, and the claimed broad state-of-the-art performance across AudioSet and multiple downstream tasks would be of considerable interest to the audio SSL community. However, as submitted, the manuscript provides no evidence whatsoever for these claims. There is no method section, no mathematical formulation of the MCL integration, no experimental setup, no tables, no error bars, and no comparison to prior work. The only substantive content is an unrelated PDE-estimation paper. Consequently, the significance of the claimed contribution cannot be assessed, and the manuscript in its present form has no scientific content relevant to its stated topic.","major_comments":[{"comment":"The document is internally inconsistent at the most basic level. The abstract describes a self-supervised audio representation learning method, MATPAC++, with MCL and evaluation on AudioSet and downstream audio tasks. The full text is a manuscript about estimating PDE and delay-PDE models from data, authored by different authors (Mai, Kroll, Thiele, Kamps), with no mention of MATPAC, MCL, masked latent prediction, audio, or any of the claimed experiments. The central claim of state-of-the-art performance is therefore completely unsupported: no architecture, equations, training protocol, or results for MATPAC++ appear anywhere in the submitted text. This is a load-bearing failure that cannot be repaired by local revision.","section":"Abstract vs. Full Text"},{"comment":"Even granting the abstract's content, the submitted text contains no description of the proposed MCL-based prediction or classification pretext tasks. There is no formulation of the predictor module, no explanation of how multiple choice hypotheses are generated or selected, no loss function, and no architectural details. Without these, the claimed contribution is neither reproducible nor checkable, and the abstract's assertions are unsupported assertions rather than scientific claims.","section":"Absence of method description"},{"comment":"The abstract claims state-of-the-art fine-tuning results on AudioSet and overall state-of-the-art downstream scores, as well as improved efficiency for music-only training. The submitted text contains no experimental section, no dataset descriptions, no metrics, no comparison baselines, and no results tables. There is no way to verify these empirical claims or to assess whether the alleged improvements come from MCL or from other unstated factors such as model capacity or training schedule. The claimed experimental superiority is entirely unsubstantiated.","section":"Absence of empirical evaluation"}],"minor_comments":[{"comment":"The title of the submitted full text, 'Hyperparameter Optimization in the Estimation of PDE and Delay-PDE models from data', matches arXiv:2508.12715, not the MATPAC++ title in the submission metadata. The abstract and body are from entirely different works; this should be resolved at the submission level before any review can proceed.","section":"Title and metadata"},{"comment":"The bibliography and appendix of the submitted text are those of the PDE paper and are irrelevant to the abstract's claims. The text also contains typographical issues (e.g., 'sensitivtiy', 'spars', 'accomodate') that would need correction in any eventual revision of the intended manuscript.","section":"References and formatting"}],"recommendation":"reject","confidential_remarks":"This is not a case of a debatable modeling choice or insufficient experimental support; the submitted manuscript file is simply a different paper from the one described in the abstract. The editor should confirm that the file was correctly uploaded. If the authors intended to submit a MATPAC++ paper, that paper is not present; the PDE manuscript, on its own, may be a legitimate work but it does not belong under this submission identifier. Rejection is the only feasible recommendation for the current submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the thing: the abstract and metadata describe MATPAC++, an audio self-supervised learning paper, but the body is a completely different paper on PDE estimation. That means the submission is not evaluable as-is. None of the claimed contributions—MCL integration, AudioSet fine-tuning results, music-domain efficiency—appear anywhere in the text. There are no equations, no experimental setup, no tables. This isn't a subtle modeling weakness; it's a total absence of evidence for the central claims.\n\nI want to be fair: the abstract's idea—using Multiple Choice Learning to model ambiguity in masked prediction—is a plausible extension of existing SSL work. And the PDE paper that occupies the body looks like a legitimate piece of work on Bayesian optimization for PDE estimation, with benchmarks on Allen-Cahn and reaction-diffusion systems. But that paper is arXiv:2508.12715, not 12709. Whatever the cause, the mismatch is load-bearing: a reviewer can't check a single claim in the abstract.\n\nThe reader's low-confidence reject is right. I'd go further: this should be desk rejected, not sent to peer review, because there is no content to review. If the authors intended to submit a different paper, they can resubmit the correct MATPAC++ manuscript. If this is a packaging error, it needs to be fixed before anything else. I have no basis to judge the actual MATPAC++ method from this submission; the abstract alone isn't enough. So the verdict is: don't spend referee time on this. Send it back.","headline":"Submission mismatched: abstract claims a MATPAC++ audio SSL paper, body is an unrelated PDE paper—no support for any of the central claims.","tokens_in":8161,"tokens_out":1685,"would_cite":false,"duration_ms":19507,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MATPAC++ claims that integrating Multiple Choice Learning into the masked prediction and unsupervised classification pretext tasks of MATPAC yields state-of-the-art self-supervised audio representations, with top AudioSet fine-tuning and do","keywords":["self-supervised learning","audio representation learning","masked latent prediction","multiple choice learning","prediction ambiguity","AudioSet","linear probing","music pretraining"],"falsifier":"Compare MATPAC++ against MATPAC under identical pretraining data, compute, and protocols; if linear-probe or AudioSet fine-tuning scores are not higher, the central claim fails. Also, in the supplied PDF, the body does not contain the MATPAC++ experiments at all, so inspecting the manuscript itself already falsifies the claim that the full text supports the abstract.","tokens_in":7448,"feed_emoji":"🎵","tokens_out":4466,"duration_ms":53185,"temperature":0.7,"pith_summary":"The paper proposes MATPAC++, an enhancement of the MATPAC self-supervised audio model, arguing that the predictor at the output of a masked latent prediction system should explicitly handle the ambiguity of audio containing multiple simultaneous sources. To do this, it integrates Multiple Choice Learning (MCL) into both the prediction and the unsupervised classification pretext tasks, so the model must propose several candidate answers rather than one. The authors report that this integration yields state-of-the-art results when fine-tuned on AudioSet and overall state-of-the-art downstream scores, and that a music-only variant reaches state-of-the-art with better efficiency. A caveat, located in the submitted full text: the body and appendices are actually a different manuscript on estimating partial differential equations from data, so the only verifiable content here is the abstract; the empirical claims cannot currently be checked.","feed_headline":"Ambiguity-aware prediction lifts audio SSL to top scores","feed_subtitle":"MATPAC++ trains masked predictors to propose several answers; authors report leading AudioSet and downstream results.","key_machinery":"Multiple Choice Learning (MCL): a training scheme in which a prediction module outputs $K$ candidate hypotheses and the loss is evaluated against the best candidate (often with an oracle/assignment step), so different hypotheses specialize to different modes of the target distribution. In MATPAC++ it is inserted into the two pretext tasks — masked latent prediction and unsupervised classification — replacing the single-target predictor of MATPAC. Its job is to model the inherent ambiguity of audio with multiple overlapping sources, and it is the load-bearing change claimed to improve representation quality.","core_discovery":"On its own terms, the contribution is a specific architectural and training change: MATPAC++ keeps MATPAC's masked latent prediction framework but replaces deterministic prediction with Multiple Choice Learning, where the predictor generates several hypotheses and the training loss is computed against the best-matching hypothesis (or a subset). The same MCL treatment is applied to the unsupervised classification pretext task. The intended effect is that the encoder can no longer collapse the multiple plausible continuations of masked audio into a single averaged target; instead it learns representations that are informative about the ambiguity itself. The authors claim this gives better tran","pith_inferences":["If the benefit of MCL comes primarily from having multiple hypotheses rather than from the best-loss selection, then a simpler multi-head predictor sharing the same capacity might reproduce part of the gain; this is testable by ablating the oracle assignment.","The ambiguity argument predicts the largest gains on polyphonic examples with several simultaneous sources; downstream tasks dominated by single-source sounds should show smaller improvements, which could be checked by stratifying AudioSet classes by source count.","Because the submitted full text is a different paper, an immediate next step is locating the actual MATPAC++ experimental section; until then the abstract-level claims rest on the authors' reporting alone."],"forward_implications":["If the claim holds, making the pretext predictor explicitly multi-modal is sufficient to raise the quality of self-supervised audio representations without changing the backbone or dataset.","AudioSet fine-tuning becomes a benchmark where MATPAC++ places above prior state-of-the-art SSL audio models under the paper's unified evaluation protocol.","Linear probing on downstream tasks improves, indicating the learned representations transfer better to tasks beyond pretraining.","Music-only pretraining with MCL gives state-of-the-art performance with substantially better efficiency, suggesting ambiguity modeling matters especially in music data.","The unified protocol enables comparisons across methods, so previously reported gaps between SSL audio methods may need re-measuring under one protocol."],"supporting_citations":[],"fun_headline_variants":["Multiple choice audio SSL tops state of the art","Handling audio ambiguity with choice learning lifts SSL","MATPAC++: ambiguity-aware SSL reaches audio SOTA","Predicting multiple futures turns audio SSL to SOTA","Choice learning resolves audio ambiguity, boosting SSL"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that explicit multiple-choice conditioning on ambiguous audio — not added capacity, a different backbone, or extra training budget — is what improves the learned representations, and that the abstract's description matches the submitted experiments (the full text currently does not).","fun_headline_variants_meta":{"raw":{"variants":["Multiple choice audio SSL tops state of the art","Handling audio ambiguity with choice learning lifts SSL","MATPAC++: ambiguity-aware SSL reaches audio SOTA","Predicting multiple futures turns audio SSL to SOTA","Choice learning resolves audio ambiguity, boosting SSL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001106,"raw_usage":{"total_tokens":4434,"prompt_tokens":714,"completion_tokens":3720,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":3661}},"tokens_in":458,"tokens_out":3720,"duration_ms":31437,"temperature":1.0,"reasoning_tokens":3661,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:17:38.114474+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare MATPAC++ against MATPAC under identical pretraining data, compute, and protocols; if linear-probe or AudioSet fine-tuning scores are not higher, the central claim fails. Also, in the supplied PDF, the body does not contain the MATPAC++ experiments at all, so inspecting the manuscript itself already falsifies the claim that the full text supports the abstract.","supporting_citations":[],"review_version":1}