{"id":"9cb5bdbe-af03-4220-90ef-f79b72ea7da9","arxiv_id":"2501.16171","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A source separation model that takes hyperellipsoid region queries in an embedding space, allowing users to control both the target location and its spread, reports strong performance on MoisesDB.","lead":"This paper presents a music source separation system that can extract any combination of instruments a user specifies as an 'ellipsoid' in an abstract sound-embedding space, with adjustable broadness. It aims to move beyond the standard four-stem (vocals, drums, bass, others) paradigm, which could give musicians a more flexible remixing tool if the reported results hold up.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-source SOTA claim is supported only by clip-wise oracle choice of query-scale α, making the reported SNR an upper bound rather than achievable performance.","rationale":"I considered two candidate concerns: (1) whether PaSST embeddings are discriminative enough for hyperellipsoid queries to be meaningful, and (2) whether the reported performance numbers support the SOTA claim. The reader's weakest_assumption was (1), but the more load-bearing issue is (2), because it is a direct threat to the paper's headline even under perfect embeddings. The single-source evaluation explicitly uses 'the clip-wise best α' (Section 4.1, Figure 5). Since α is the query broadness parameter, selecting it per clip from the test target is an oracle procedure. This is not a minor implementation detail; it changes the evaluation from 'what can a user achieve without knowing the answer' to 'what is the best possible query.' Combined with the full-track versus clip-wise protocol difference against Banquet and the absence of any fixed-stem SOTA baseline, the 'state-of-the-art' claim is unsupported as written. The idea itself remains plausible and the multi-source retrieval metrics are useful, so the appropriate verdict is CONDITIONAL: accept only after a non-oracle, matched-protocol evaluation is provided. I do not see an internal inconsistency in the hyperellipsoid formalism; the concern is about evidence, not construction.","tokens_in":17201,"tokens_out":3566,"duration_ms":34719,"concrete_test":"Recompute the single-source evaluation using one fixed α per stem chosen only from the validation set (e.g., the α maximizing median validation SNR), then evaluate the test clips with that fixed α and no per-clip selection. If the resulting median SNRs fall below the Banquet numbers reported in Table B.II for stems such as Grand Pf or Org, then the claimed SNR advantage is an artifact of the clip-wise oracle. As a second check, rerun Banquet on the identical clip-wise, 10-s sliding-window protocol with the proposed system's fixed-α results; if Banquet's clip-wise numbers meet or exceed the proposed system's, the state-of-the-art claim fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 4.1 states that for each clip the model is queried with radii r̃ = αr⊥ for α ∈ [10^-3, 1], and Figure 5 reports 'the clip-wise best α' as the proposed method's result. Because α controls the broadness of the hyperellipsoidal query, choosing it per clip using the ground-truth target gives the system an oracle over its own query parameter. A user of the system would have to fix α before hearing the output, so the median SNRs in Figure 5 and Table B.II are an upper envelope, not a deployable system's expected performance. The comparison is further mismatched: Banquet was evaluated full-track with overlap-add while the proposed method is evaluated clip-wise, so the 'on par or better' conclusion conflates oracle selection with protocol differences. The abstract's 'state-of-the-art performance ... in terms of signal-to-noise ratios' therefore rests on an invalid comparison. The retrieval metrics, being ROC-style curves over α, are less affected by this issue, but they do not rescue the SNR claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a music source separation system that generalizes query-by-example to query-by-region: the target is specified by a hyperellipsoid in the embedding space of a pretrained PaSST audio classifier, with the ellipsoid center and semi-axes controlling which instrument embeddings are included. This enables arbitrary composite targets and user-controllable broadness. The model is a FiLM-conditioned encoder-decoder trained on MoisesDB with queries precomputed from all source subsets, using enclosing/excluding ellipsoids to define valid targets. The paper reports single-source and multi-source SNR and retrieval metrics, claiming state-of-the-art performance.","tokens_in":17389,"tokens_out":4960,"duration_ms":45813,"significance":"If the results are taken at face value, the system would be a significant step beyond fixed-stem separation: a single model can extract arbitrary composites specified geometrically, including long-tail instruments for which prior query-based systems collapsed. The retrieval-evaluation methodology via least-squares projection is a useful contribution. The use of a pretrained, frozen PaSST embedding avoids the circularity of training the query space itself, and the training-data generation from all subsets is a systematic approach. However, the headline SOTA claim is not yet supported by the evaluation protocol.","major_comments":[{"comment":"The single-source SNR figures are computed using the clip-wise best α, i.e., an oracle selection of the query scale factor per clip. Since α is a query parameter a user must set before hearing the output, these median SNRs are an upper envelope over α rather than the performance of any deployable configuration. The abstract's claim of state-of-the-art SNR therefore rests on an oracle evaluation. Please report performance at a fixed α, or average over α with a specified selection rule, and discuss the trade-off between broadness and SNR.","section":"Section 4.1, Figure 5, Table B.II"},{"comment":"The only comparative baseline, Banquet, is evaluated under a different protocol: full-track overlap-add for Banquet versus clip-wise evaluation for the proposed system. The paper itself cautions that the comparison is only a rough gauge, yet the abstract and conclusion claim state-of-the-art performance. With a single mismatched baseline, this claim is not supported. Please add matched-protocol comparisons against at least one fixed-stem SOTA system (e.g., HTDemucs) on the same MoisesDB split, and either full-track overlap-add evaluation for the proposed system or clip-wise evaluation for the baseline.","section":"Section 4.1, Table B.II"},{"comment":"The query precomputation restricts training and evaluation to target subsets for which an enclosing hyperellipsoid excludes all non-target sources; if a non-target embedding falls inside the enclosing ellipsoid, that source is removed from the mixture. This guarantees that every query is separable in the PaSST space by construction, so the evaluation does not measure how often the query-by-region formalism fails for realistic subsets. The paper should report the fraction of subsets discarded or made infeasible, and should evaluate on all subsets (including non-separable ones) to test the underlying assumption that PaSST embeddings cluster by instrument.","section":"Section 3.1"},{"comment":"The thresholded retrieval metrics (accuracy, precision, recall, F1) require a decision threshold on the least-squares scores, and Table B.III reports a threshold per stem without stating how it was selected. If these thresholds are tuned on the test set, the metrics are optimistic. Please state the selection procedure (e.g., validation-set optimization or a fixed threshold), or report unthresholded metrics such as ROCAUC/PRAUC as primary.","section":"Section 3.2, Table B.III"}],"minor_comments":[{"comment":"The vector representation q^T = [c^T tril(K)^T] is underspecified: tril(K) is not defined in the text, and the dimension of q should be stated explicitly. Please clarify the ordering of the lower-triangular entries.","section":"Section 2.1, Eq. (8)"},{"comment":"The notation switches between Λ = diag(λ) and r = λ^{1/2} in the definitions of inclusion/exclusion radii; please use a single consistent notation to avoid confusion.","section":"Section 3.1"},{"comment":"The caption appears truncated after the hyperplane definition; the sentence about the two-dimensional cross-section is incomplete and should be rewritten for clarity.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable extension of the authors' prior Banquet work; the novelty lies in the hyperellipsoid query representation and the retrieval evaluation. The main concern is the evaluation protocol: the oracle α selection and the single mismatched baseline do not support the state-of-the-art claim in the abstract. I would recommend that the editor require the authors to provide fixed-α results, matched baseline comparisons, and a report of query feasibility rates before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is genuine: generalizing Banquet's point queries to hyperellipsoidal region queries gives users control over target broadness, and the least-squares projection method for retrieval evaluation is a practical response to a real evaluation gap. The geometry is clearly formalized, and the precomputation of valid query intervals via enclosing and excluding ellipsoids is a sensible way to generate training data. I also credit the authors for explicitly noting the clip-wise versus full-track evaluation mismatch in Figure 5, even though the abstract overstates what that comparison supports.\n\nThe main soft spot is the evaluation protocol. In Section 4.1, the model is queried with a sweep of alpha values per clip and the clip-wise best alpha is reported. That is an oracle over the query-width parameter, not something a user can reproduce without knowing the ground-truth target in advance. The median SNRs in Figure 5 and Table B.II are therefore an upper envelope, not the expected performance of a deployable system. This directly undermines the abstract's \"state-of-the-art performance\" claim for SNR. The comparison to Banquet is also the only baseline, and it is mismatched in protocol; a fair comparison would need at least one non-self baseline evaluated under identical conditions.\n\nThe retrieval metrics are less affected by the oracle-alpha issue, but they have their own caveats: the least-squares projection weights can be unstable when sources are highly correlated, and the per-stem thresholds appear to be chosen on the test data, which risks optimism. Still, as a relative diagnostic the retrieval evaluation is informative—it shows, for example, that the system is recall-heavy, which makes sense for a region-query method.\n\nThe underlying assumption that PaSST embeddings are discriminative enough for ellipsoidal regions to correspond to musically meaningful target sets is plausible but not tested. A quick analysis of embedding separability for the MoisesDB stems would help, especially because the training-query precomputation only keeps subsets that are ellipsoid-separable; evaluation under the same protocol may be more optimistic than a user's arbitrary query would be.\n\nThis is a competent, internally consistent paper with a novel query geometry and a useful evaluation idea. The performance claims need to be fixed—no oracle alpha, fair baselines, and ideally released code or more training details. I would send it to peer review, but with major revision expected.","headline":"The hyperellipsoid query is a real extension of the authors' Banquet system, and the retrieval evaluation is a useful addition, but the state-of-the-art SNR claim rests on an oracle choice of the query-width parameter and a mismatched baseline.","tokens_in":17954,"tokens_out":1536,"would_cite":true,"duration_ms":16909,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Music source separation can be driven by hyperellipsoidal region queries: a single model extracts any stem or composite whose embedding lies inside a user-drawn region, with controllable broadness, and reports state-of-the-art results on…","keywords":["music source separation","query-by-region","hyperellipsoid query","stem-agnostic separation","MoisesDB","PaSST embedding","retrieval metrics","FiLM conditioning"],"falsifier":"Compute, for every stem in MoisesDB, the smallest enclosing hyperellipsoid around that stem's embeddings and count how often a non-target stem's embeddings fall inside; if pairs such as kick drum and bass guitar, which share low-frequency energy, cannot be isolated by any ellipsoid, per-pair retrieval should fall to chance, which would refute the claim that region queries can specify arbitrary targets.","tokens_in":16949,"feed_emoji":"🎧","tokens_out":14586,"duration_ms":115505,"temperature":0.7,"pith_summary":"Music source separation has long been locked to a fixed menu of stems—vocals, drums, bass, and \"other\"—because the best models are trained to output exactly those four. This paper tries to break that lock: it claims a single model can extract any target, a lone instrument or an arbitrary blend of instruments, when the target is specified as a hyperellipsoid (a multidimensional ellipse) drawn in a pretrained audio-embedding space. The user sets both where the ellipse sits (the timbre wanted) and how wide it is (how much of the surrounding timbral neighborhood to include), and the model pulls out exactly the sources whose embeddings fall inside it. If the claim holds, one trained model replaces stacks of single-stem extractors, reaches instrument classes it never saw in training, and gives musicians a single continuous knob for extraction broadness. The paper reports state-of-the-art signal-to-noise ratios and retrieval scores on MoisesDB in support of that claim.","feed_headline":"Draw an ellipse; one model separates whatever lies inside it","feed_subtitle":"No fixed stem lists: one model extracts any instrument or composite, with user-controlled broadness.","key_machinery":"The load-bearing object is the hyperellipsoid query $Q(c,K) = \\{z \\in \\mathbb{R}^P : (z-c)^\\top K^{-1}(z-c) \\leq 1\\}$, a Mahalanobis-distance ball specified by a center $c$ and a positive-definite spread matrix $K$. The query space is the 768-dimensional PaSST embedding, reduced by PCA to 128 dimensions (91.8\\% explained variance on the training set), and the query is packed into a vector of the center plus the lower-triangular entries of $K$, which a small fully connected network maps to FiLM parameters $\\gamma, \\beta$ that rescale and shift the mixture embedding at the decoder bottleneck. Training queries are precomputed per clip by finding, for every possible target subset, the smallest ellipsoid enclosing the target embeddings and the largest same-center ellipsoid excluding all non-target embeddings, then sampling radii uniformly between the two. A level-matching regularizer with adaptive weighting keeps the output from collapsing to near silence, which the authors report happens without it.","core_discovery":"The central claim is that query-by-region works as a general formalism for music source separation: given a mixture and a hyperellipsoid in a discriminative embedding space, the model recovers exactly the sum of the sources whose embeddings lie inside the ellipsoid, regardless of how many sources or which classes the target set contains. The paper extends the point-query Banquet architecture so that the conditioning input is a full hyperellipsoid—center plus positive-definite spread matrix—mapped to FiLM parameters that adapt the mixture embedding at the bottleneck of a time-frequency masking network. Because a hyperellipsoid is the level set of a Mahalanobis distance, it is a natural geometric stand-in for a multivariate Gaussian cluster, and interpolating between a smallest enclosing ellipsoid and a largest excluding ellipsoid generates valid training queries for every source subset in each clip. On MoisesDB the system reports state-of-the-art SNR and retrieval metrics, including macro and micro average precision of 0.83 and 0.86, and it recovers long-tail instruments (organ, synth, brass, reeds, strings) where its point-query predecessor collapsed to silence.","pith_inferences":["Editorial inference: the hyperellipsoid is a level set of a multivariate Gaussian, so the natural next step—one the paper lists as future work—is a Gaussian-mixture query that extracts sources from several disjoint timbral regions at once; the same FiLM conditioning machinery should carry over unchanged.","Editorial inference: the method's ceiling is set by the embedding space, not the decoder; swapping the PaSST query space for another pretrained or task-fine-tuned embedding and checking whether mAP tracks instrument-discriminability would isolate where the query-by-region gains come from.","Editorial inference: the class-dependent sensitivity to query width suggests that optimal broadness is a property of the target's timbral neighborhood; an automatic radius selector tuned on validation retrieval metrics would remove the need to sweep the scale factor at test time."],"forward_implications":["A single trained model can extract any single stem or any composite target whose sources can be enclosed by a hyperellipsoid, including instrument classes never seen in training: viola is absent from the training set yet is extracted at a median SNR of 6.1 dB.","Users gain a continuous broadness control: scaling the query radii toward the excluding ellipsoid widens the extraction, and the reported ROC analysis shows the effect is class-dependent, with bass guitar insensitive to the scale factor while grand piano and brass degrade markedly at the wrong setting.","Long-tail instruments that collapsed in the point-query predecessor are recovered: organs, synths, brass, reeds, and strings all move from zero SNR to positive median SNR, with the largest gains on exactly the classes that the fixed-stem paradigm serves worst.","The least-squares projection evaluation turns an audio-format separation output into per-source retrieval scores, giving query-based systems a way to separate \"did it find the right sources\" from \"is the audio clean,\" and yields macro and micro average precision of 0.83 and 0.86.","Performance tracks the fraction of the mixture requested: median SNR and weighted mean average precision both rise as the target-to-mixture source ratio grows, so query difficulty behaves like a standard retrieval setting where the relevant proportion of the collection sets the difficulty."],"supporting_citations":[{"why":"The point-query Banquet model this work extends; supplies the decoder architecture, the multi-domain multichannel L1SNR loss, the data-split convention, and the baseline for single-source comparisons.","marker":"[Watcharasupat and Lerch, 2024]"},{"why":"PaSST, the pretrained audio transformer whose 768-dimensional embedding space is the query space; the discriminability of this space is the assumption that makes ellipsoid queries meaningful.","marker":"[Koutini et al., 2022]"},{"why":"MoisesDB, the multi-stem dataset that provides the sources, mixtures, and evaluation splits for all experiments.","marker":"[Pereira et al., 2023]"},{"why":"The earlier region-query system in a low-dimensional hyperbolic space; its fidelity limit motivates using a higher-dimensional embedding with hyperellipsoid regions.","marker":"[Petermann et al., 2023]"},{"why":"HTDemucs, the fixed-stem state of the art that recent query-based systems, including this one, position themselves against.","marker":"[Rouard et al., 2023]"},{"why":"Demucs, the source of the global input normalization and denormalization adopted to stabilize training.","marker":"[Défossez et al., 2019a]"}],"fun_headline_variants":["One model separates any stem with a hyperellipse query","Draw a hyperellipse to isolate any set of instruments","Hyperellipsoidal queries enable stem-agnostic separation","Separation by region: ellipse any target, no fixed stems"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pretrained PaSST audio-embedding space groups sounds by instrument well enough that a hyperellipsoid drawn in it always isolates a musically meaningful target set from the non-target sounds in the same mixture.","fun_headline_variants_meta":{"raw":{"variants":["One model separates any stem with a hyperellipse query","Draw a hyperellipse to isolate any set of instruments","Hyperellipsoidal queries enable stem-agnostic separation","Separation by region: ellipse any target, no fixed stems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1639,"prompt_tokens":1067,"completion_tokens":572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":504}},"tokens_in":683,"tokens_out":572,"duration_ms":5853,"temperature":1.0,"reasoning_tokens":504,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:40:01.606474+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute, for every stem in MoisesDB, the smallest enclosing hyperellipsoid around that stem's embeddings and count how often a non-target stem's embeddings fall inside; if pairs such as kick drum and bass guitar, which share low-frequency energy, cannot be isolated by any ellipsoid, per-pair retrieval should fall to chance, which would refute the claim that region queries can specify arbitrary targets.","supporting_citations":[{"cited_title":"Efficient Training of Audio Transformers with Patchout","cited_arxiv_id":null,"evidence_quote":"PaSST, the pretrained audio transformer whose 768-dimensional embedding space is the query space; the discriminability of this space is the assumption that makes ellipsoid queries meaningful."},{"cited_title":"MoisesDB : A Dataset for Source Separation Beyond 4- Stems","cited_arxiv_id":null,"evidence_quote":"MoisesDB, the multi-stem dataset that provides the sources, mixtures, and evaluation splits for all experiments."},{"cited_title":"Hyperbolic Audio Source Separation","cited_arxiv_id":null,"evidence_quote":"The earlier region-query system in a low-dimensional hyperbolic space; its fidelity limit motivates using a higher-dimensional embedding with hyperellipsoid regions."},{"cited_title":"Hybrid Transformers for Music Source Separation","cited_arxiv_id":null,"evidence_quote":"HTDemucs, the fixed-stem state of the art that recent query-based systems, including this one, position themselves against."}],"review_version":1}